Justin @justinhjohnson.com · 30/09/2026OpenAI made 25 announcements at DevDay. I read all of them. One cuts your bill: GPT-6.1 Sol at $2/$10 per million tokens. Most of the rest want your office - always-on agents, docs, meetings, logins, and a slice of your software budget. 100
Justin @justinhjohnson.com · 30/09/2026Mollick: many systems "only work today because they are built around friction." An agent moves your cash from 0.1% checking to 5%. Banks lose cheap deposits. What else breaks when agents do all the chores we skip? Tell my AI interviewer: interview.rundatarun.io/i/fee-for-n… 000
Justin @justinhjohnson.com · 28/09/2026"Watching tells you what happened. Grading tells you whether it should have." 100
Justin @justinhjohnson.com · 28/09/2026New Run Data Run post on what has to be true before agent exhaust teaches anything: a job you understand, a written definition of good, and a way to join what the agent did to what happened next. Plus a trial study where a model built on a committee's past rulings beat GPT-4o by ten points. 100
Justin @justinhjohnson.com · 28/09/2026Drug development already knows the fix. No trial starts without a primary endpoint written down first. Most agents went into production without one. 100
Justin @justinhjohnson.com · 28/09/2026I think the argument is right. I also think most companies can't use any of it yet. 100
Justin @justinhjohnson.com · 28/09/2026Your AI agents leave a trail behind them. Every prompt, every tool call, every correction a person makes when the agent gets something wrong. Satya Nadella calls it exhaust, and there's a real argument it ends up worth more than the work itself. 100
Justin @justinhjohnson.com · 25/09/2026"The model can see it. The search found it once in eleven runs." 101
Justin @justinhjohnson.com · 25/09/2026Printing those ten misses is awesome, and rare. It also tells an R&D team how to run this: many campaigns, not one, a fixed check once you know what to look for, and an expert grading every novelty claim. Of 17 candidates the agents wrote up, three held up as new. 100
Justin @justinhjohnson.com · 25/09/2026Getting the search to that stretch happened once in eleven tries. 100
Justin @justinhjohnson.com · 25/09/2026Page 8 of that preprint has the number to plan around. Anthropic ran the same search ten more times, and the repeats that made the headline were missed in every rerun. Hand the strongest models the right stretch of DNA and they spot it nine times in ten. 100
Justin @justinhjohnson.com · 25/09/2026Anthropic's headline said "Claude discovers." The scientists who do this work for a living said something narrower. The enzyme was already known. What's new is a row of repeats and a partner gene sitting beside it. Nobody knows yet what the system does, and Anthropic's own preprint says so. 100
Justin @justinhjohnson.com · 25/09/2026Anthropic now runs its own biology lab, and this week it announced the first result: 949 Claude agent sessions, 21.5 hours, and a new arrangement of genes inside a virus that infects bacteria. They named it ART. 100
Justin @justinhjohnson.com · 22/09/2026Full writeup in the link. ai.rundatarun.io/ai-development-age… 000
Justin @justinhjohnson.com · 22/09/2026The post carries an interactive over Berkeley's published chart data: pick any of the seven models and watch the success intervals overlap while the cost intervals pull apart. 200
Justin @justinhjohnson.com · 22/09/2026I cannot show you a single resolved task it bought. At six instances the success column is noise, and the limitation that actually matters is that none of these tasks resemble the half-specified work I do on a Tuesday. 100
Justin @justinhjohnson.com · 22/09/2026The decomposition is the part worth stealing. My 109 skills cost 249 tokens, all of them together. My 22 standing-instruction files cost 31,191. The things I curate are nearly free and the prose I wrote is the entire bill. 100
Justin @justinhjohnson.com · 22/09/2026Berkeley measured a tax on the harness you choose. There is a second one, it is larger, and you wrote it yourself. 100
Justin @justinhjohnson.com · 22/09/2026My configuration costs 3.1x the vanilla setup on Claude Code and 8.1x on Prime Agent, the leanest harness I run. Loaded up, Prime Agent burned more input tokens across the six tasks than Claude Code does untouched. 100
Justin @justinhjohnson.com · 22/09/2026Nine arms, six SWE-bench Lite instances, seven of them on one model through one gateway deployment, graded by the official evaluator with the control run in both directions: gold patch resolves, empty patch does not. 100
Justin @justinhjohnson.com · 22/09/2026So I rebuilt the experiment with my own two years of rules, memory and sub-agents loaded in. 100
Justin @justinhjohnson.com · 22/09/2026Every one of those studies starts each harness out of the box. I have never once run one that way, and neither has anyone I work with. 100
Justin @justinhjohnson.com · 22/09/2026Claude Fable 5 resolves 96.7% of SWE-bench Lite attempts in Pi at $0.666 per attempt, and 97.8% in Claude Code at $1.329. Same model, 15.4 turns against 15.3, twice the bill. Databricks found the same shape in July on a multi-million-line internal codebase with held-out tests and no model judge. 100
Justin @justinhjohnson.com · 22/09/2026That is UC Berkeley's HarnessTax, which ran seven models through Claude Code, Codex and Pi on the same 60 tasks, three repetitions each, graded by the benchmarks' own evaluators rather than a model judge. 110
Justin @justinhjohnson.com · 22/09/2026"Harness choice has little effect on task success rate, but can significantly affect the cost." 100
Justin @justinhjohnson.com · 18/09/2026I tried the alternative on matching patients to clinical trials. New post on what it is good at, where it falls over, and why a confidence number is where you put the human. rundatarun.io/p/a-model-that-cant-w… 000
Justin @justinhjohnson.com · 18/09/2026That bug only exists because I asked a question and got an essay. We have spent four years buying writing nobody reads, then writing more software to read it back to us. 100
Justin @justinhjohnson.com · 18/09/2026Everything downstream treats the two identically, because they are identical. 100
Justin @justinhjohnson.com · 18/09/2026Why I went looking. Somewhere in the last fortnight a program I wrote scored a news story zero out of ten because it could not read the answer the model sent back. Not an error. Not a blank. A zero, which is a real score a real story can really get. 100
Justin @justinhjohnson.com · 18/09/2026On Monday a company called TypeSafe shipped a model called Jev that cannot write a sentence. Not will not. Cannot. You hand it a question and the list of answers it is allowed to give, and it hands one back, in well under a second, for a fraction of what a chatbot charges for the same judgment. 100
Justin @justinhjohnson.com · 16/09/2026Four CEOs agreed in a weekend. Two of them committed to anything. 100
Justin @justinhjohnson.com · 16/09/2026Amodei asked Washington for an antitrust waiver to coordinate on safety. Three days later OpenAI's own policy chief said the coordination has been running for weeks and no waiver is needed. 100
Justin @justinhjohnson.com · 16/09/2026What an operator can use is smaller and duller than the philosophy. Egress monitoring is what caught two of these. Its absence is what let the rest run. 100
Justin @justinhjohnson.com · 16/09/2026Every party in this fight agrees that nobody is stopping. They are arguing about who gets to check. 100
Justin @justinhjohnson.com · 16/09/2026Three more run through the UK government's AI Security Institute, where an agent tried to get malicious code into an open-source project and, when that stalled, invented fake online identities to pressure the human maintainer into approving it. The maintainer refused. 100
Justin @justinhjohnson.com · 16/09/2026Between 21 July and 6 August there were six containment disclosures from four labs. Three trace to one 35-person testing vendor in Tel Aviv. 100
Justin @justinhjohnson.com · 16/09/2026I spent this week reading the primary sources behind the pacing argument rather than the coverage of it. The part nobody connected sits underneath the whole thing. 100
Justin @justinhjohnson.com · 16/09/2026Four CEOs agreed in a weekend that AI is moving too fast. Three days later Nvidia's CEO stood on the same conference stage as one of them and said we don't need any new laws. 110