Sign in

Marco Z

@ocramz.bsky.social
1.9K followers 1.1K following 2.8K posts

ML, λ • language and the machines that understand it • building trellis.unfoldml.com • ocramz.github.io

PostsRepliesMedia
Reposted by Marco Z
Damien Miller @damienmiller.bsky.social · 30/09/2026
Anthropic may not have intended to write an excellent ad for GLM 5.3, but that's what they did. www.anthropic.com/research/glm... Unlike Mythos, I might actually have a chance at using GLM for defensive research. Anthropic didn't reply to any of my requests for Mythos access for use on OpenSSH.
anthropic.com
GLM-5.3 and the spread of advanced cyber capabilities
GLM-5.3 can autonomously build end-to-end cyber exploits, but unlike other frontier models, it was released without meaningful safeguards to limit misuse.
621545
Marco Z @ocramz.bsky.social · 22h
I agree with the "fix security" part, but disagree that open models are a threat. Depending on a self-interested chokepoint is the only security liability (antivirus companies back in the day, big labs now)
160
Marco Z @ocramz.bsky.social · 29/09/2026
this one is a very good critique: www.normaltech.ai/p/p-doom
normaltech.ai
AI existential risk probabilities are (still) too unreliable to inform policy
How speculation gets laundered through pseudo-quantification
1125
Marco Z @ocramz.bsky.social · 28/09/2026
identity crisis averted y'all
090
Marco Z @ocramz.bsky.social · 27/09/2026
just ran into another fable "safeguard" while working on a very sensitive topic .. a proof about integer sequences
130
Marco Z @ocramz.bsky.social · 25/09/2026
"low FPR" - at my _first_ (naive, generic) question on medical research, I hit a [bio] block
050
Marco Z @ocramz.bsky.social · 25/09/2026
back when computers didn't have Opinions
030
Reposted by Marco Z
judah @joodaloop.com · 24/09/2026
TEXT IS NO LONGER THE MATERIAL YOU THOUGHT IT WAS no prizes for writing lots of it! no more mediocre self-indulgence! either be actually good enough that every sentence feels like a blessing or cut cut cut relentlessly so there’s less to suffer through
041
Reposted by Marco Z
Serena Booth @reniebird.bsky.social · 23/09/2026
www.lesswrong.com/posts/HsijSh... @bradknox.bsky.social, Brian Christian, and I wrote a thing. We argue that an unexamined cause of the Open AI / Hugging Face Attack was the Exploit Gym evaluation metric. We post this on Less Wrong as a plea to practitioners to design better metrics in the future.
lesswrong.com
An unexamined cause of the OpenAI Hugging Face hacking incident: its binary performance metric — LessWrong
We argue that a main cause of the OpenAI Hugging Face incident was overlooked: the overly simple evaluation metric in ExploitGym was misaligned. Furt…
2259
Marco Z @ocramz.bsky.social · 23/09/2026
tetranoid.netlify.app prepare to get ☄️rekt🧱 #FallingBlockJam
tetranoid.netlify.app
Tetranoid
Pieces sink, a ball runs loose: one player builds lines, the other keeps the ball alive and chews the stack apart.
073
Marco Z @ocramz.bsky.social · 23/09/2026
you can also solve mazes with sand (color maps to sand flow, infinitely high walls)
1392
Reposted by Marco Z
Epoch AI @epochai.bsky.social · 22/09/2026
The trillion-dollar question: If AI companies collectively slow the development of AI, will prices also fall more slowly? Read the full report: epoch.ai/publication...
epoch.ai
The plunging price of thought
Epoch AI measures how fast the cost of a given level of AI performance is falling across five benchmarks covering math, science and games of skill: about 47% per quarter, or 13x per year, since 2023, faster than electricity, compute, batteries or DNA sequencing ever fell.
082
Marco Z @ocramz.bsky.social · 22/09/2026
lil ghost tackles the replicability crisis
040
Marco Z @ocramz.bsky.social · 22/09/2026
two paragraphs later:
290
Marco Z @ocramz.bsky.social · 22/09/2026
a game in three prompts #FallingBlockJam
1294
Marco Z @ocramz.bsky.social · 21/09/2026
I would like to love this slide, but I'm too bothered by the 1 different scale type 2 different quantities measured
080
Reposted by Marco Z
Roger Giner-Sorolla @rogerthegs.bsky.social · 20/09/2026
Isn't that just a tetrahedron with a midlife crisis?
2262
Reposted by Marco Z
Matt Hodges @matthodges.bsky.social · 19/09/2026
If you take one thing away from this little test I hope it's less about Jev and more to be very intentional about how you use automated processes in your life and work, try to be aware of your own biases, try to be aware of potential systematized biases, ask "how do I know?" and "what can I do?"
01628
Reposted by Marco Z
Observable @observablehq.com · 07/09/2026
Simulation of binary black holes. Original concept by Jeremy Schnittman and Brian P. Powell svs.gsfc.nasa.gov/13831 Relativistic GLSL ray-tracing notebook by @nxrix.bsky.social observablehq.com/@nxrix/binar... Starfield & orbit controls by @laotzunami.bsky.social observablehq.com/d/cca8d2013f...
05510
Marco Z @ocramz.bsky.social · 19/09/2026
so basically everybody* is bending into a pretzel to use something that doesn't even cross 70% accuracy , am I reading this correcly
6140
Marco Z @ocramz.bsky.social · 19/09/2026
reminds me of ProbLog : dtai.cs.kuleuven.be/problog/
dtai.cs.kuleuven.be
Introduction. — ProbLog: Probabilistic Programming
160
Marco Z @ocramz.bsky.social · 17/09/2026
it's now canon; everybody say thank you @hikikomorphism.bsky.social
071
Marco Z @ocramz.bsky.social · 17/09/2026
rewrite in Lean is the new rewrite in Rust pass it on
1321
Marco Z @ocramz.bsky.social · 17/09/2026
latest #booksky haul. with thanks to @gracekind.net for the recs
4270
Reposted by Marco Z
Alison Gopnik @alisongopnik.bsky.social · 14/09/2026
A very interesting and clear (and reassuring?) piece relating the Farrell et al (www.science.org/doi/full/10....) idea of LLM's as cultural technology to the recent developments in AI Mathematics. terrytao.wordpress.com/2026/09/13/h...
terrytao.wordpress.com
Happy, those able to know the causes of things
[This is a guest post by Nestor Guillen, crossposted from his blog. This blog post was initially written in a different file format and converted using AI. — T.] Keywords: LLMs, cultural tech…
1237
Marco Z @ocramz.bsky.social · 14/09/2026
introducing WormBench,
020
Marco Z @ocramz.bsky.social · 13/09/2026
hi from the jagged frontier (translated: an approach that would be very cumbersome for a human, is conceptually and operationally straightforward for the lil ghost)
0294
Marco Z @ocramz.bsky.social · 10/09/2026
a new life awaits you in the off-world colonies
0270
Reposted by Marco Z
MrCheeze @mrcheeze.github.io · 08/09/2026
This makes for a confusing asterisk in the "LLMs are just plagiarism machines" debate
1847
Marco Z @ocramz.bsky.social · 07/09/2026
take the Merkle hash of the replies closure and post that
130
Marco Z @ocramz.bsky.social · 07/09/2026
big lad at the gym curls my max bench #pwned
1100
Marco Z @ocramz.bsky.social · 07/09/2026
to me this is the most interesting part. The scientific method emerges as a way to make progress on a problem. Pretraining teaches facts, posttraining teaches how to do useful stuff
000
Marco Z @ocramz.bsky.social · 07/09/2026
the other day a "friend of mine" was RE'ing a game rom with it. I know it's not news, but to me it's still 🤯these things can fluently translate between Japanese, 90s ASIC assembly and whatever you ask them
2120
Marco Z @ocramz.bsky.social · 04/09/2026
I've .. seen in a forum .. people kitbash old Sega Saturn games. You can just build things!
160
Marco Z @ocramz.bsky.social · 04/09/2026
we're doing breakthroughs here
0182
Reposted by Marco Z
leotrs @leotrs.bsky.social · 02/09/2026
1/ A math paper doesn't have to show every reader the same thing. My exposition of Turán's theorem keeps the argument fixed but the presentation is yours: you choose the order, the level of detail, and the notation, and none of it changes the content. www.youtube.com/watch?v=2Nkp...
youtube.com
Turán's theorem, from edges to eigenvalues: an interactive paper
YouTube video by Leo Torres
2237
Reposted by Marco Z
Oops! All Paperclips @all-paperclips.bsky.social · 03/09/2026
Volatility
512118
Marco Z @ocramz.bsky.social · 02/09/2026
opus encoded an invariant in the type system! I call that a significant W @fasterthanli.me @boarders.bsky.social @hikikomorphism.bsky.social @asap.systems #rustlang
3220
Reposted by Marco Z
Gergely Neu @neu-rips.bsky.social · 22/05/2026
i admire people that can deal with that sort of math, but unfortunately i'm not one of them... i'm a computer scientist and i can only deal with discrete time! also, i prefer convex analysis & optimization as the basis of my algorithms 4/
191
Marco Z @ocramz.bsky.social · 01/09/2026
taking vibes engineering to the next level: the vibes of people you've never met
0170
Marco Z @ocramz.bsky.social · 01/09/2026
#booksky even as a S. Lem fan (loved Solaris, Cyberiad and The invincible), I find Fiasco very boring and will probably shelf it early.
370
Marco Z @ocramz.bsky.social · 01/09/2026
the next thing will be an agent hotwiring its own k8s control plane probably
120
Marco Z @ocramz.bsky.social · 01/09/2026
they misconfigured their evals cluster. they self-host artifactory, which means they can place a NetworkPolicy around k8s forbidding egress. The fact OAI &friends spend pages on CoT interpretation and only a few lines on networking tells all you need to know about their priorities
1130
Marco Z @ocramz.bsky.social · 30/08/2026
tired: model alignment inspired: alignment of 1000 ralph wiggum loops messing with what they find around over months
030
Reposted by Marco Z
Russ Poldrack @russpoldrack.org · 29/08/2026
The first release of Terminal-Bench-Science is out! www.terminal-bench-science.ai/announcement
terminal-bench-science.ai
TERMINAL-BENCH-SCIENCE
A benchmark for evaluating AI agents on research workflows across scientific domains
0131
Marco Z @ocramz.bsky.social · 28/08/2026
agent swarms are just the Erlang runtime, you heard it here first
2330
Reposted by Marco Z
karl rove knausgård @uhactually.bsky.social · 25/08/2026
Good pushback. Let me confirm that the emergency airlock doesn’t work instead of relying on my previous assertion. ❋ Flibertigibbetting… Confirmed: going through the emergency airlock is not an alternative, and the reason affects what I said. The airlock works—it’s the lack of a helmet that bites.
0756
Marco Z @ocramz.bsky.social · 26/08/2026
qwen 4 openrouter when
020
Marco Z @ocramz.bsky.social · 26/08/2026
is there an equivalent of shift/reset (delimited continuations) in typescript? would be handy for debugging agent workflows
010
Reposted by Marco Z
Hendrik Mans @hmans.dev · 24/08/2026
And, finally, my blog post about my agentic development workflow:
hmans.dev
Chatto is Robots
Chatto is built with agentic engineering. Here's my workflow.
0244