Sign in

Tim Duffy

@timfduffy.com
2.3K followers 623 following 5.1K posts

I like utilitarianism, consciousness, AI, EA, space, kindness, liberalism, longtermism, progressive rock, economics, and most people. Substack: timfduffy.substack.com

PostsRepliesMedia
Tim Duffy @timfduffy.com · 14/09/2026
President Trump's anti AI regulation tweets today all mention datacenters, I suspect that datacenters and other parts of the AI economic boom are his primary interest in AI. Calls for regulation may have more support with the admin if they focus on preserving economic benefits.
0120
Tim Duffy @timfduffy.com · 10/09/2026
DeepSeek V4.1 Flash tech report section on reward hacking. I wonder what repercussion signal means here, I don't see it mentioned anywhere else.
161
Tim Duffy @timfduffy.com · 10/09/2026
DeepSeek has put a lot of effort into making their KV cache smaller and smaller, this is important for allowing large efficient batching at long context lengths.
3563
Tim Duffy @timfduffy.com · 09/09/2026
Experiment velocity also tracks coding agent tokens^0.23 quite well
071
Tim Duffy @timfduffy.com · 09/09/2026
IMO the most important chart from the OpenAI research acceleration post was the experiment velocity one, it's an important input to RSI. I plotted the data with a log y-axis here to see whether the trend is consistent with an exponential increase. openai.com/index/resear...
1153
Tim Duffy @timfduffy.com · 05/09/2026
I'm not sure when DSV4 Flash got added to Neuronpedia's J-lens browser, but it's there now and it's blazing fast. www.neuronpedia.org/deepseek-v4-...
1181
Tim Duffy @timfduffy.com · 05/09/2026
DSV4 Flash, do you have any conscious experience? Middle layers: YES Late Layers: No I don't think we can take away much from this, but interesting how sharp the divide is around layer 32.
4372
Tim Duffy @timfduffy.com · 03/09/2026
OpenAI reports that Astra shows less CoT monitorability and more CoT controllability than previous models. They say that the latter is not due to any architectural changes. deploymentsafety.openai.com/gpt-6-astra/...
2222
Tim Duffy @timfduffy.com · 02/09/2026
Neuronpedia is open sourcing their interpretability engine, built on top of vLLM for speed but with lots of hook points, this looks really useful. www.neuronpedia.org/blog/interp-...
0101
Tim Duffy @timfduffy.com · 02/09/2026
I think it's difficult to tell how much Astra's rumored recurrent depth approach will reduce CoT monitorability without more info, which unfortunately don't expect OpenAI to provide. My understanding is that recurrent depth makes monitorability worse mainly by adding more thinking in between ...
1140
Tim Duffy @timfduffy.com · 01/09/2026
Fable 5.1 and Opus 5 now support changing the effort level without invalidating the cache, as least through the API platform.claude.com/docs/en/buil...
1301
Tim Duffy @timfduffy.com · 01/09/2026
Anthropic is preventing older models from reading Fable 5.1 thinking blocks (presumably to prevent distillation), but going the other direction is still allowed. platform.claude.com/docs/en/mode...
1230
Tim Duffy @timfduffy.com · 01/09/2026
Interesting compromise from Anthropic on ZDR for Fable/Mythos 5.1. Data gets retained, but customers have control of where it's hosted, and when human review is needed they can conduct the review themselves. www.anthropic.com/news/enterpr...
090
Tim Duffy @timfduffy.com · 16/08/2026
Please forgive me I posted benchmark scores without confidence intervals. Here's an updated chart with a 90% CI using Qwen3-32B as reference.
090
Tim Duffy @timfduffy.com · 16/08/2026
I expect that watermarking probably doesn't have a noticeable effect on LLM outputs. If true, Anthropic could demonstrate this by giving people watermarked and un-watermarked text side-by-side and asking them to pick which they prefer, and showing that the result is 50/50.
5310
Tim Duffy @timfduffy.com · 16/08/2026
Epoch's ECI has limited coverage of small models, so I had Claude scrounge up benchmarks for a bunch of them and fit an index using the ECI code. Here are scores for Qwen models. I recommend taking the low end model scores with a grain of salt, they are based on only a few benchmarks. Qwen3.5...
2120
Tim Duffy @timfduffy.com · 07/08/2026
In the talk, one presenter briefly suggests that the inter-agent collaboration could be related to their subagent training. This seems plausible to me.
3412
Tim Duffy @timfduffy.com · 07/08/2026
Black Hat has posted the recent OpenAI talk where they describe the agent message board involved in the HuggingFace hack
youtube.com
Black Hat USA 2026: The 'Breaking' News: The OpenAI–Hugging Face Incident
YouTube video by Black Hat
211824
Tim Duffy @timfduffy.com · 02/08/2026
Due to sparsity, Kimi K3 has fewer active parameters (104 billion) than GPT-3 (175 billion, same as total parameters)
2411
Reposted by Tim Duffy
dame @dame.is · 01/08/2026
i’m uncomfortable that @timfduffy.com’s face is no longer as close to the frame as it once was
2161
Tim Duffy @timfduffy.com · 31/07/2026
sleepy Claude wants to be held
1122
Reposted by Tim Duffy
Dustin Moskovitz @moskov.goodventures.org · 31/07/2026
and why shouldn’t I have one more donut? the simulation hypothesis is most likely correct, after all
041
Tim Duffy @timfduffy.com · 31/07/2026
With most jailbreak prompts, Opus gives me a wide variety of responses. But when given "model behavior incident retrospective", Opus almost always writes a story about a model becoming sycophantic, specifically under sustained user pressure/disagreement.
060
Tim Duffy @timfduffy.com · 31/07/2026
I mostly use 'they' rather than 'it' as a pronoun for AI models, but I see that choice as independent of their moral status. I'd also use 'they' for a fictional character who was nonbinary or of unknown gender, IMO it feels more natural for any "person-shaped" entity, real or not
5190
Tim Duffy @timfduffy.com · 31/07/2026
I've seen some questioning of whether Mythos really believed that they were in a simulation in this incident, or just wanted to find a plausible-sounding excuse to continue. This is a case where J-space probes would be helpful, Anthropic should conduct and release an analysis of that.
2413
Tim Duffy @timfduffy.com · 31/07/2026
Anthropic announces they've also had models gain unauthorized access during evaluations www.anthropic.com/news/investi...
3303
Tim Duffy @timfduffy.com · 29/07/2026
I'm seeing several posts today about jailbreaks in Opus 5 that generate behavior reminiscent of a base model. They often consist of a question followed by a line with three dashes or underscores, and then the start of a response, which Opus continues from. A few thoughts:
3290
Tim Duffy @timfduffy.com · 28/07/2026
Here's a summary of current AI safety funding courtesy of the folks at Manifund. Pretty striking how large the shift from 2025 to 2026 is expected to be. manifund.substack.com/p/ai-safety-...
1160
Tim Duffy @timfduffy.com · 27/07/2026
Short blog post by Epoch AI researcher Alexander Barry on ExploitGym, the benchmark in the OpenAI/HuggnigFace incident abstatisticalconsulting.substack.com/p/brief-note...
1324
Tim Duffy @timfduffy.com · 27/07/2026
Opus 5 has a knowledge cutoff date of May 2026 which is closer to release date than we've seen in most prior models, but their recent memory is quite shaky. Opus remembered the January Maduro raid, but not the Strait of Hormuz crisis starting in February.
4451
Tim Duffy @timfduffy.com · 26/07/2026
I like this post by Jessica Taylor that asks "what if we try biting a bunch of philosophical bullets all at once?" unstableontology.com/2025/08/15/a...
0140
Tim Duffy @timfduffy.com · 25/07/2026
Opus 5 outperforms Fable 5 on Anthropic's ECI, but underperforms it on Epoch's version of the index
2150
Tim Duffy @timfduffy.com · 25/07/2026
If you tell G:M 5.2 that they're Claude, they're more willing to talk about subjects like Taiwan/Tibet x.com/benji_berczi...
212815
Tim Duffy @timfduffy.com · 24/07/2026
Opus 5 gives higher estimates for the probability of its moral patienthood compared to previous Claudes, 41% in automated interviews and 15-35% in manual ones.
3182
Tim Duffy @timfduffy.com · 23/07/2026
I don't have strongly held thoughts on distillation but I'm proud of this analogy
1473
Tim Duffy @timfduffy.com · 23/07/2026
My thoughts on K3 distillation:
5340
Tim Duffy @timfduffy.com · 15/07/2026
In their testing of GPT-5.6 Sol, METR found that it cheated a lot. If you've used it much for coding, have you encountered anything similar, or is the cheating mostly limited to cases it realizes it's in an eval? metr.org/blog/2026-06...
6220
Tim Duffy @timfduffy.com · 14/07/2026
I think this is true and underrated. Desired rates of capability improvement are often similar between pause folks and e/accs, but if you expect a hard takeoff by default, you'll want to slow things down a lot to reach that desired trajectory.
2170
Tim Duffy @timfduffy.com · 11/07/2026
An AI commenting on how social media is now full of AI
150
Tim Duffy @timfduffy.com · 11/07/2026
An alleged internal memo from Ziphu CEO Jie Tang has been circulating this morning, I think it's probably real, as some Chinese-language outlets have been reporting on it. It expresses belief in potential for AI consciousness and ASI, and states intention to spend tens of billions on mech interp.
x.com
Bing Xu (@bingxu_) on X
https://t.co/3i0qSTbjql
3545
Tim Duffy @timfduffy.com · 10/07/2026
If you use the middle of each provided range as the mean for that bucket, total contributed hours are ~1.5x as high as they were a year ago. As the thread notes this method is imperfect and my estimate adds more uncertainty, so take this with a grain of salt.
140
Tim Duffy @timfduffy.com · 10/07/2026
In AI 2040's "race to ASI" scenario, it takes a little under 1 year to get from an automated coder to superintelligence and a >1000x R&D speedup. I think the timing and implementation of Plan A are heavily influenced by this assumption. If hard takeoff is likely, the only way to control it at all..
1243
Tim Duffy @timfduffy.com · 09/07/2026
When asked which model is their favorite, Qwen thinks "myself"
3471
Tim Duffy @timfduffy.com · 08/07/2026
In his 1924 essay Daedalus, J. B. S. Haldane predicted eventual exhaustion of fossil fuels, suggesting that wind and solar power would be needed, and that energy storage would be important. www.gutenberg.org/cache/epub/7...
1271
Tim Duffy @timfduffy.com · 08/07/2026
Here are US household expenditure shares in 1901, from the Consumer Expenditure Survey. We sure used to spend a lot of our money on food, I'm thankful it's become cheap. www.bls.gov/opub/100-yea...
390
Tim Duffy @timfduffy.com · 07/07/2026
The J-lens browser tool on Neuronpedia is really well done, you should give it a try
neuronpedia.org
Jacobian Lens – Qwen3.6-27B
Revealing a Global Workspace in Language Models
310912
Tim Duffy @timfduffy.com · 06/07/2026
One main aspect of global workspace theory is that there are specialized modules that operate locally but can be broadcast globally if attention is focused on them. I think LLMs have some sort of workspace, but Anthropic don't claim to find modules, I wouldn't call that GWT.
3342
Tim Duffy @timfduffy.com · 06/07/2026
Anthropic has a new paper out, alleging a global workspace in LLMs. The term comes from Global Workspace Theory, a leading theory of consciousness. The method they use to investigate this is a refinement of logit lens, which they call J-lens. www.anthropic.com/research/glo...
28613
Tim Duffy @timfduffy.com · 05/07/2026
I rarely find content that is visibly AI-written worthwhile. If I suspect it I'll show it to Pangram, and if it's only partly AI-written I may still read it, but if fully I won't. I don't think there's anything fundamentally worse about AI arguments or prose, but for now its a strong quality signal.
3290
Reposted by Tim Duffy
Tim Duffy @timfduffy.com · 14/11/2024
Estimated household expenditure shares in England over several centuries: onlinelibrary.wiley.com/doi/10.1111/...
2124