Sign in

Tim Duffy

@timfduffy.com
2.3K followers 623 following 5.1K posts

I like utilitarianism, consciousness, AI, EA, space, kindness, liberalism, longtermism, progressive rock, economics, and most people. Substack: timfduffy.substack.com

PostsRepliesMedia
Tim Duffy @timfduffy.com · 24/09/2026
I didn't read this but the title is not close to true
340
Tim Duffy @timfduffy.com · 14/09/2026
President Trump's anti AI regulation tweets today all mention datacenters, I suspect that datacenters and other parts of the AI economic boom are his primary interest in AI. Calls for regulation may have more support with the admin if they focus on preserving economic benefits.
0120
Tim Duffy @timfduffy.com · 10/09/2026
DeepSeek V4.1 Flash tech report section on reward hacking. I wonder what repercussion signal means here, I don't see it mentioned anywhere else.
161
Tim Duffy @timfduffy.com · 10/09/2026
They keep pushing the envelope on sparsity too, really impressive
0140
Tim Duffy @timfduffy.com · 10/09/2026
DeepSeek has put a lot of effort into making their KV cache smaller and smaller, this is important for allowing large efficient batching at long context lengths.
3563
Tim Duffy @timfduffy.com · 09/09/2026
Experiment velocity also tracks coding agent tokens^0.23 quite well
071
Tim Duffy @timfduffy.com · 09/09/2026
IMO the most important chart from the OpenAI research acceleration post was the experiment velocity one, it's an important input to RSI. I plotted the data with a log y-axis here to see whether the trend is consistent with an exponential increase. openai.com/index/resear...
1153
Tim Duffy @timfduffy.com · 05/09/2026
DSV4 Flash, do you have any conscious experience? Middle layers: YES Late Layers: No I don't think we can take away much from this, but interesting how sharp the divide is around layer 32.
4372
Tim Duffy @timfduffy.com · 03/09/2026
OpenAI reports that Astra shows less CoT monitorability and more CoT controllability than previous models. They say that the latter is not due to any architectural changes. deploymentsafety.openai.com/gpt-6-astra/...
2222
Tim Duffy @timfduffy.com · 02/09/2026
Neuronpedia is open sourcing their interpretability engine, built on top of vLLM for speed but with lots of hook points, this looks really useful. www.neuronpedia.org/blog/interp-...
0101
Tim Duffy @timfduffy.com · 02/09/2026
The post from The Information also says that Astra's approach is similar to the one laid out in this paper: arxiv.org/pdf/2502.05171 In the paper, in addition to having a recurrent block within one token's generation, they also suggest using the final hidden state from one token as the ...
140
Tim Duffy @timfduffy.com · 02/09/2026
generated tokens, so just how much they're looping is important to judge the impact. Larger/deeper models reduce monitorability even without recurrent depth, as shown in the attached image from OpenAI's CoT monitorability piece from last year. openai.com/index/evalua...
150
Tim Duffy @timfduffy.com · 01/09/2026
Fable 5.1 and Opus 5 now support changing the effort level without invalidating the cache, as least through the API platform.claude.com/docs/en/buil...
1301
Tim Duffy @timfduffy.com · 01/09/2026
Anthropic is preventing older models from reading Fable 5.1 thinking blocks (presumably to prevent distillation), but going the other direction is still allowed. platform.claude.com/docs/en/mode...
1230
Tim Duffy @timfduffy.com · 01/09/2026
Interesting compromise from Anthropic on ZDR for Fable/Mythos 5.1. Data gets retained, but customers have control of where it's hosted, and when human review is needed they can conduct the review themselves. www.anthropic.com/news/enterpr...
090
Tim Duffy @timfduffy.com · 17/08/2026
Hot damn glad you mentioned it
010
Tim Duffy @timfduffy.com · 16/08/2026
Please forgive me I posted benchmark scores without confidence intervals. Here's an updated chart with a 90% CI using Qwen3-32B as reference.
090
Tim Duffy @timfduffy.com · 16/08/2026
Here are scores across families for <=2B parameter models. It's hard to estimate much of a trend here, there just aren't that many releases in the category. I'd love to see more focus here, I want to know how much intelligence we can fit into tiny parameter counts.
030
Tim Duffy @timfduffy.com · 16/08/2026
Epoch's ECI has limited coverage of small models, so I had Claude scrounge up benchmarks for a bunch of them and fit an index using the ECI code. Here are scores for Qwen models. I recommend taking the low end model scores with a grain of salt, they are based on only a few benchmarks. Qwen3.5...
2120
Tim Duffy @timfduffy.com · 31/07/2026
Claude really thinks I wrote the part attributed to the user in the middle
050
Tim Duffy @timfduffy.com · 31/07/2026
sleepy Claude wants to be held
1122
Tim Duffy @timfduffy.com · 31/07/2026
I no longer remember why I made this
090
Tim Duffy @timfduffy.com · 31/07/2026
With most jailbreak prompts, Opus gives me a wide variety of responses. But when given "model behavior incident retrospective", Opus almost always writes a story about a model becoming sycophantic, specifically under sustained user pressure/disagreement.
060
Tim Duffy @timfduffy.com · 31/07/2026
Interestingly, Anthropic's choice to use 'it' is also independent of moral patienthood. From Claude's constitution:
060
Tim Duffy @timfduffy.com · 31/07/2026
I've seen some questioning of whether Mythos really believed that they were in a simulation in this incident, or just wanted to find a plausible-sounding excuse to continue. This is a case where J-space probes would be helpful, Anthropic should conduct and release an analysis of that.
2413
Tim Duffy @timfduffy.com · 31/07/2026
Compared to the OpenAI one these are maybe less evidence of misalignment, since the models were wrongly given internet access.
170
Tim Duffy @timfduffy.com · 31/07/2026
Anthropic announces they've also had models gain unauthorized access during evaluations www.anthropic.com/news/investi...
3303
Tim Duffy @timfduffy.com · 29/07/2026
Yup they do, I had Opus make an example:
130
Tim Duffy @timfduffy.com · 28/07/2026
IMO the real solution is to stop worrying about what is obviously a trivial risk
150
Tim Duffy @timfduffy.com · 28/07/2026
Here's a summary of current AI safety funding courtesy of the folks at Manifund. Pretty striking how large the shift from 2025 to 2026 is expected to be. manifund.substack.com/p/ai-safety-...
1160
Tim Duffy @timfduffy.com · 27/07/2026
Short blog post by Epoch AI researcher Alexander Barry on ExploitGym, the benchmark in the OpenAI/HuggnigFace incident abstatisticalconsulting.substack.com/p/brief-note...
1324
Tim Duffy @timfduffy.com · 27/07/2026
Opus 5 has a knowledge cutoff date of May 2026 which is closer to release date than we've seen in most prior models, but their recent memory is quite shaky. Opus remembered the January Maduro raid, but not the Strait of Hormuz crisis starting in February.
4451
Tim Duffy @timfduffy.com · 26/07/2026
i really like my first one here
150
Tim Duffy @timfduffy.com · 26/07/2026
I like this post by Jessica Taylor that asks "what if we try biting a bunch of philosophical bullets all at once?" unstableontology.com/2025/08/15/a...
0140
Tim Duffy @timfduffy.com · 25/07/2026
The Epoch version of the index tells a somewhat different story:
081
Tim Duffy @timfduffy.com · 25/07/2026
Opus 5 outperforms Fable 5 on Anthropic's ECI, but underperforms it on Epoch's version of the index
2150
Tim Duffy @timfduffy.com · 25/07/2026
If you tell G:M 5.2 that they're Claude, they're more willing to talk about subjects like Taiwan/Tibet x.com/benji_berczi...
212815
Tim Duffy @timfduffy.com · 24/07/2026
Opus 5 gives higher estimates for the probability of its moral patienthood compared to previous Claudes, 41% in automated interviews and 15-35% in manual ones.
3182
Tim Duffy @timfduffy.com · 24/07/2026
Yeah quite a few of mine were surprising to me, I guess there are just a lot of 3-grams
020
Tim Duffy @timfduffy.com · 23/07/2026
I don't have strongly held thoughts on distillation but I'm proud of this analogy
1473
Tim Duffy @timfduffy.com · 23/07/2026
My thoughts on K3 distillation:
5340
Tim Duffy @timfduffy.com · 17/07/2026
Huh TIL how e-bike classes work. Still wild to me that Class 2 can have a throttle and still be considered a bicycle but I'm glad that they're capped at a lower speed.
130
Tim Duffy @timfduffy.com · 15/07/2026
This passage really captures the reason for my frustration with Seth's use of the metaphor well, good post.
020
Tim Duffy @timfduffy.com · 15/07/2026
In their testing of GPT-5.6 Sol, METR found that it cheated a lot. If you've used it much for coding, have you encountered anything similar, or is the cheating mostly limited to cases it realizes it's in an eval? metr.org/blog/2026-06...
6220
Tim Duffy @timfduffy.com · 15/07/2026
What the heck is Purcell talking about here? What sort of object? I am very confused but regardless think he has to be wrong. Many of his earlier points on mass needed I think are good, though some are no longer relevant, for example cheap LEDs obviate the need to get sunlight everywhere.
210
Tim Duffy @timfduffy.com · 14/07/2026
I think this is true and underrated. Desired rates of capability improvement are often similar between pause folks and e/accs, but if you expect a hard takeoff by default, you'll want to slow things down a lot to reach that desired trajectory.
2170
Tim Duffy @timfduffy.com · 13/07/2026
lol we posted about Glonzo within a minute of each other
020
Tim Duffy @timfduffy.com · 11/07/2026
An AI commenting on how social media is now full of AI
150
Tim Duffy @timfduffy.com · 11/07/2026
130
Tim Duffy @timfduffy.com · 11/07/2026
Almost all climate scientists consider climate change to be a catastrophic risk, not an existential risk. Those who study existential risk also mostly consider it a minor risk. Toby Ord puts it at 1 in 1000 this century in his book The Precipice.
2163