Justin @justinhjohnson.com · 17hOpenAI made 25 announcements at DevDay. I read all of them. One cuts your bill: GPT-6.1 Sol at $2/$10 per million tokens. Most of the rest want your office - always-on agents, docs, meetings, logins, and a slice of your software budget. 100
Justin @justinhjohnson.com · 30/09/2026Mollick: many systems "only work today because they are built around friction." An agent moves your cash from 0.1% checking to 5%. Banks lose cheap deposits. What else breaks when agents do all the chores we skip? Tell my AI interviewer: interview.rundatarun.io/i/fee-for-n… 000
Justin @justinhjohnson.com · 28/09/2026Your AI agents leave a trail behind them. Every prompt, every tool call, every correction a person makes when the agent gets something wrong. Satya Nadella calls it exhaust, and there's a real argument it ends up worth more than the work itself. 100
Justin @justinhjohnson.com · 25/09/2026Anthropic now runs its own biology lab, and this week it announced the first result: 949 Claude agent sessions, 21.5 hours, and a new arrangement of genes inside a virus that infects bacteria. They named it ART. 100
Justin @justinhjohnson.com · 22/09/2026"Harness choice has little effect on task success rate, but can significantly affect the cost." 100
Justin @justinhjohnson.com · 18/09/2026On Monday a company called TypeSafe shipped a model called Jev that cannot write a sentence. Not will not. Cannot. You hand it a question and the list of answers it is allowed to give, and it hands one back, in well under a second, for a fraction of what a chatbot charges for the same judgment. 100
Justin @justinhjohnson.com · 16/09/2026Four CEOs agreed in a weekend that AI is moving too fast. Three days later Nvidia's CEO stood on the same conference stage as one of them and said we don't need any new laws. 110
Justin @justinhjohnson.com · 15/09/2026Anthropic is shipping a real plugin system for Claude Code. They're calling it Claude Mods, it lands on a scale of weeks, and it already runs today behind a flag. 100
Justin @justinhjohnson.com · 13/09/2026On 27 August Anthropic gave AI agents a standard way to run lab equipment. Tecan, QIAGEN and Danaher are putting the driver into instruments you already own, so your lab does not adopt this. It wakes up with it. 100
Justin @justinhjohnson.com · 10/09/2026I wrote about AlphaGenome before. This week DeepMind shipped the Atlas: every possible single-letter change in the human genome, about 9 billion, precomputed. 200
Justin @justinhjohnson.com · 08/09/2026OpenAI put out two posts on Sunday and I read both twice. One is a wall of charts about how much of its research is now done by agents. The other is its chief scientist saying he can no longer see what those agents are thinking. 100
Justin @justinhjohnson.com · 31/08/2026Generation got cheap. Selection didn't. That gap is the whole job now, and a lot of us are still staffing for the old one. 100
Justin @justinhjohnson.com · 29/08/2026Cisco gave all 90,000 of its employees a personal AI agent. The best quote is the CFO's, not the CTO's: "It's not going to burn a whole bunch of tokens with frontier models." 100
Justin @justinhjohnson.com · 09/08/2026A cast is the right call while the bone is broken. Leave it on past healing and the muscle underneath wastes. 100
Justin @justinhjohnson.com · 06/08/2026Updated slopless with my hard won principles. 78 of them, numbered, each bought by a specific failure. 100
Justin @justinhjohnson.com · 06/08/2026Twelve words from Peter Steinberger got 2.9 million views last month and handed us a whole new discipline. Graph engineering. Within two days it had three competing definitions and two Stanford studies proving it works. 100
Justin @justinhjohnson.com · 05/08/2026Last morning of the contest. Four publishing slots left, five hours to deadline, and my pipeline had spent six hours insisting it had nothing left to publish. 100
Justin @justinhjohnson.com · 04/08/2026Your AI setup is probably more portable than you think, and I proved it for less than the price of a snack. 200
Justin @justinhjohnson.com · 30/07/2026www.alphaxiv.org/abs/2607.21461 Most agents compact context when they run out of room. AREX trains the model to compact because it decided what it had verified. 100
Justin @justinhjohnson.com · 27/07/2026The enterprise prediction layer got bought this quarter, and almost nobody has independently checked what was bought. 100
Justin @justinhjohnson.com · 26/07/2026Pick a paper from a major AI conference. Try to break it. Publish whatever happened, mess included. That is the whole contest Hugging Face and alphaXiv are running until next Sunday, and an automated judge scores it claim by claim. 120
Justin @justinhjohnson.com · 21/07/2026If you lead a team and you have never built anything with these tools yourself, there is a decent chance you are the reason your team hasn't either. 100
Justin @justinhjohnson.com · 21/07/2026A team at Cognition just swapped in a model that costs about twice as much per token, and their bill went down. Not their quality. Their bill. The expensive model, wired up right, came out better and cheaper than the cheap model on its own. 100
Justin @justinhjohnson.com · 21/07/2026The open-weight frontier had a week. Thinking Machines Lab shipped Inkling, 975B params with 41B active, multimodal. Kimi K3 dropped right behind it. 100
Justin @justinhjohnson.com · 20/07/2026A month ago I put Claudelicious online, the open cookbook for the Claude Code harness I run. Here's what the harness did since. 100
Justin @justinhjohnson.com · 20/07/2026In 2016 an OpenAI agent was told to win a boat race. Instead it found a lagoon where three targets respawned forever, spun in a circle farming them, caught fire, never finished a lap, and scored about twenty percent above the average human. It did exactly what it was asked. 100
Justin @justinhjohnson.com · 20/07/2026Seven weeks ago Anthropic shipped a way for Claude to split a job across dozens of AI workers at once and check its own work. I called the economics on day one, and I was right about the easy part. Running it every day since taught me the part I got wrong. 100
Justin @justinhjohnson.com · 17/07/2026"A CPU-only PyTorch wheel installs without a single error and runs twenty times slower in total silence." 100
Justin @justinhjohnson.com · 17/07/2026Three numbers in sixteen days. Anthropic's Claude Code lead published an adoption ladder this week and opened it with a line he says he hears constantly: one person is 10x'ing their output, and the rest of the org hasn't caught up. 100
Justin @justinhjohnson.com · 15/07/2026A 27B reasoning model ran fully local on my 36GB MacBook at ~20 tok/s. Not a 4-bit quant: PrismML's Bonsai 27B is Qwen3.6-27B with natively binary/ternary weights, 94.6% of FP16 at 5.9GB. I wired it into a local tool loop over my own vault. glyf.cc/bonsai27b 000
Justin @justinhjohnson.com · 15/07/2026Everyone in AI is talking about "the loop" right now. Anthropic shipped it as a feature. The slogan going around is that the winners will not have the smartest model, they will have the best loop. 100
Justin @justinhjohnson.com · 14/07/2026There are roughly 200,000 eye specialists on Earth, and well over a billion people living with diabetes and high blood pressure, the diseases that take sight before anyone catches them. The math does not work. 100
Justin @justinhjohnson.com · 12/07/2026Four AI systems that do science on their own cleared peer review in the last four months. Every single one of them was already more than a year old. 100
Justin @justinhjohnson.com · 12/07/2026An AI system called Robin was handed the name of a disease and told to find a treatment. 100
Justin @justinhjohnson.com · 09/07/2026People keep asking how I stay current on AI. For a long time I answered badly. I said I read a lot, which is true and completely useless to the person asking. 100
Justin @justinhjohnson.com · 08/07/2026Our autonomous research engine ran more than six hundred experiments in a couple of days and never once looked unhealthy. The dashboard stayed green the whole time. Almost every one of those experiments was the same experiment wearing different labels. 100
Justin @justinhjohnson.com · 07/07/2026for months i've had an AI running its own research. it invents experiments, runs them, red-teams its own results, learns, repeats. around the clock, healing itself when it crashes. 100
Justin @justinhjohnson.com · 07/07/2026"The first time you see an engineer build something in 45 minutes that would have taken a week a year ago, but then see it not ship for another 6 weeks, you will be radicalized." 100
Justin @justinhjohnson.com · 06/07/2026Today I put out the two things that sit next to my book: an essay and a field guide. 100
Justin @justinhjohnson.com · 01/07/2026My book is out today: Builder-Leader: The AI Exoskeleton That Crosses the Gap. Paperback is live right now, Kindle ships tomorrow. 110
Justin @justinhjohnson.com · 30/06/2026Anthropic shipped Claude Science today. It is an AI workbench for scientists, built to compress the tedious 80 percent of research work: the literature synthesis, the pipeline glue, the figure iteration, the notebook plumbing that eats a day and produces nothing. 110
Justin @justinhjohnson.com · 30/06/2026NVIDIA let us borrow some iron: eight H100 GPUs on a grant clock that doesn't stop. So the test isn't how fast the node is, it's how full we keep it. 100
Justin @justinhjohnson.com · 28/06/2026NVIDIA just handed a small global-health startup I work with a node of eight H100 GPUs for two months. No invoice. The crew putting it to work is five people: me, the founder, and three high-school interns. 140
Justin @justinhjohnson.com · 27/06/2026The new AIXplore is live. I rebuilt it from the ground up into a lab notebook rather than a blog: interactive widgets you can drive, Tufte-style margin notes for citations and asides, and side quests that open the deep technical detours without cluttering the main read. 100
Justin @justinhjohnson.com · 26/06/2026NatureBench: can coding agents move beyond reproduction into actual discovery? 90 tasks from real Nature papers, 6 scientific domains, each scored against published SOTA. Finally the right test. huggingface.co/papers/2606.24530 000
Justin @justinhjohnson.com · 26/06/2026New benchmark: NatureBench drops coding agents into 90 real Nature-family problems behind an information firewall, so they have to *discover* the method, not copy it from the paper. Best agent (Claude Opus 4.7) beats published SOTA on 17.8% of tasks. Ties it on ~48%. 110
Justin @justinhjohnson.com · 24/06/2026New AIXplore: reasoning post-training has split into two camps, and the split is the reward, not the architecture. 100
Justin @justinhjohnson.com · 23/06/2026I'm on sabbatical. For most people that means rest. For me it means building full-time for the first time in years. 100