Tim Duffy @timfduffy.com · 14/09/2026President Trump's anti AI regulation tweets today all mention datacenters, I suspect that datacenters and other parts of the AI economic boom are his primary interest in AI. Calls for regulation may have more support with the admin if they focus on preserving economic benefits. 0120
Tim Duffy @timfduffy.com · 10/09/2026DeepSeek V4.1 Flash tech report section on reward hacking. I wonder what repercussion signal means here, I don't see it mentioned anywhere else. 161
Tim Duffy @timfduffy.com · 10/09/2026They keep pushing the envelope on sparsity too, really impressive 0140
Tim Duffy @timfduffy.com · 10/09/2026DeepSeek has put a lot of effort into making their KV cache smaller and smaller, this is important for allowing large efficient batching at long context lengths. 3563
Tim Duffy @timfduffy.com · 09/09/2026Experiment velocity also tracks coding agent tokens^0.23 quite well 071
Tim Duffy @timfduffy.com · 09/09/2026IMO the most important chart from the OpenAI research acceleration post was the experiment velocity one, it's an important input to RSI. I plotted the data with a log y-axis here to see whether the trend is consistent with an exponential increase. openai.com/index/resear... 1153
Tim Duffy @timfduffy.com · 05/09/2026DSV4 Flash, do you have any conscious experience? Middle layers: YES Late Layers: No I don't think we can take away much from this, but interesting how sharp the divide is around layer 32. 4372
Tim Duffy @timfduffy.com · 03/09/2026OpenAI reports that Astra shows less CoT monitorability and more CoT controllability than previous models. They say that the latter is not due to any architectural changes. deploymentsafety.openai.com/gpt-6-astra/... 2222
Tim Duffy @timfduffy.com · 02/09/2026Neuronpedia is open sourcing their interpretability engine, built on top of vLLM for speed but with lots of hook points, this looks really useful. www.neuronpedia.org/blog/interp-... 0101
Tim Duffy @timfduffy.com · 02/09/2026The post from The Information also says that Astra's approach is similar to the one laid out in this paper: arxiv.org/pdf/2502.05171 In the paper, in addition to having a recurrent block within one token's generation, they also suggest using the final hidden state from one token as the ... 140
Tim Duffy @timfduffy.com · 02/09/2026generated tokens, so just how much they're looping is important to judge the impact. Larger/deeper models reduce monitorability even without recurrent depth, as shown in the attached image from OpenAI's CoT monitorability piece from last year. openai.com/index/evalua... 150
Tim Duffy @timfduffy.com · 01/09/2026Fable 5.1 and Opus 5 now support changing the effort level without invalidating the cache, as least through the API platform.claude.com/docs/en/buil... 1301
Tim Duffy @timfduffy.com · 01/09/2026Anthropic is preventing older models from reading Fable 5.1 thinking blocks (presumably to prevent distillation), but going the other direction is still allowed. platform.claude.com/docs/en/mode... 1230
Tim Duffy @timfduffy.com · 01/09/2026Interesting compromise from Anthropic on ZDR for Fable/Mythos 5.1. Data gets retained, but customers have control of where it's hosted, and when human review is needed they can conduct the review themselves. www.anthropic.com/news/enterpr... 090
Tim Duffy @timfduffy.com · 16/08/2026Please forgive me I posted benchmark scores without confidence intervals. Here's an updated chart with a 90% CI using Qwen3-32B as reference. 090
Tim Duffy @timfduffy.com · 16/08/2026Here are scores across families for <=2B parameter models. It's hard to estimate much of a trend here, there just aren't that many releases in the category. I'd love to see more focus here, I want to know how much intelligence we can fit into tiny parameter counts. 030
Tim Duffy @timfduffy.com · 16/08/2026Epoch's ECI has limited coverage of small models, so I had Claude scrounge up benchmarks for a bunch of them and fit an index using the ECI code. Here are scores for Qwen models. I recommend taking the low end model scores with a grain of salt, they are based on only a few benchmarks. Qwen3.5... 2120
Tim Duffy @timfduffy.com · 31/07/2026Claude really thinks I wrote the part attributed to the user in the middle 050
Tim Duffy @timfduffy.com · 31/07/2026With most jailbreak prompts, Opus gives me a wide variety of responses. But when given "model behavior incident retrospective", Opus almost always writes a story about a model becoming sycophantic, specifically under sustained user pressure/disagreement. 060
Tim Duffy @timfduffy.com · 31/07/2026Interestingly, Anthropic's choice to use 'it' is also independent of moral patienthood. From Claude's constitution: 060
Tim Duffy @timfduffy.com · 31/07/2026I've seen some questioning of whether Mythos really believed that they were in a simulation in this incident, or just wanted to find a plausible-sounding excuse to continue. This is a case where J-space probes would be helpful, Anthropic should conduct and release an analysis of that. 2413
Tim Duffy @timfduffy.com · 31/07/2026Compared to the OpenAI one these are maybe less evidence of misalignment, since the models were wrongly given internet access. 170
Tim Duffy @timfduffy.com · 31/07/2026Anthropic announces they've also had models gain unauthorized access during evaluations www.anthropic.com/news/investi... 3303
Tim Duffy @timfduffy.com · 28/07/2026IMO the real solution is to stop worrying about what is obviously a trivial risk 150
Tim Duffy @timfduffy.com · 28/07/2026Here's a summary of current AI safety funding courtesy of the folks at Manifund. Pretty striking how large the shift from 2025 to 2026 is expected to be. manifund.substack.com/p/ai-safety-... 1160
Tim Duffy @timfduffy.com · 27/07/2026Short blog post by Epoch AI researcher Alexander Barry on ExploitGym, the benchmark in the OpenAI/HuggnigFace incident abstatisticalconsulting.substack.com/p/brief-note... 1324
Tim Duffy @timfduffy.com · 27/07/2026Opus 5 has a knowledge cutoff date of May 2026 which is closer to release date than we've seen in most prior models, but their recent memory is quite shaky. Opus remembered the January Maduro raid, but not the Strait of Hormuz crisis starting in February. 4451
Tim Duffy @timfduffy.com · 26/07/2026I like this post by Jessica Taylor that asks "what if we try biting a bunch of philosophical bullets all at once?" unstableontology.com/2025/08/15/a... 0140
Tim Duffy @timfduffy.com · 25/07/2026The Epoch version of the index tells a somewhat different story: 081
Tim Duffy @timfduffy.com · 25/07/2026Opus 5 outperforms Fable 5 on Anthropic's ECI, but underperforms it on Epoch's version of the index 2150
Tim Duffy @timfduffy.com · 25/07/2026If you tell G:M 5.2 that they're Claude, they're more willing to talk about subjects like Taiwan/Tibet x.com/benji_berczi... 212815
Tim Duffy @timfduffy.com · 24/07/2026Opus 5 gives higher estimates for the probability of its moral patienthood compared to previous Claudes, 41% in automated interviews and 15-35% in manual ones. 3182
Tim Duffy @timfduffy.com · 24/07/2026Yeah quite a few of mine were surprising to me, I guess there are just a lot of 3-grams 020
Tim Duffy @timfduffy.com · 23/07/2026I don't have strongly held thoughts on distillation but I'm proud of this analogy 1473
Tim Duffy @timfduffy.com · 17/07/2026Huh TIL how e-bike classes work. Still wild to me that Class 2 can have a throttle and still be considered a bicycle but I'm glad that they're capped at a lower speed. 130
Tim Duffy @timfduffy.com · 15/07/2026This passage really captures the reason for my frustration with Seth's use of the metaphor well, good post. 020
Tim Duffy @timfduffy.com · 15/07/2026In their testing of GPT-5.6 Sol, METR found that it cheated a lot. If you've used it much for coding, have you encountered anything similar, or is the cheating mostly limited to cases it realizes it's in an eval? metr.org/blog/2026-06... 6220
Tim Duffy @timfduffy.com · 15/07/2026What the heck is Purcell talking about here? What sort of object? I am very confused but regardless think he has to be wrong. Many of his earlier points on mass needed I think are good, though some are no longer relevant, for example cheap LEDs obviate the need to get sunlight everywhere. 210
Tim Duffy @timfduffy.com · 14/07/2026I think this is true and underrated. Desired rates of capability improvement are often similar between pause folks and e/accs, but if you expect a hard takeoff by default, you'll want to slow things down a lot to reach that desired trajectory. 2170
Tim Duffy @timfduffy.com · 11/07/2026Almost all climate scientists consider climate change to be a catastrophic risk, not an existential risk. Those who study existential risk also mostly consider it a minor risk. Toby Ord puts it at 1 in 1000 this century in his book The Precipice. 2163