Sign in

Tim Duffy

@timfduffy.com
2.3K followers 623 following 5.1K posts

I like utilitarianism, consciousness, AI, EA, space, kindness, liberalism, longtermism, progressive rock, economics, and most people. Substack: timfduffy.substack.com

PostsRepliesMedia
Tim Duffy @timfduffy.com · 24/09/2026
I didn't read this but the title is not close to true
340
Tim Duffy @timfduffy.com · 14/09/2026
President Trump's anti AI regulation tweets today all mention datacenters, I suspect that datacenters and other parts of the AI economic boom are his primary interest in AI. Calls for regulation may have more support with the admin if they focus on preserving economic benefits.
0120
Tim Duffy @timfduffy.com · 14/09/2026
A similar version I saw: "I, da AI doomer"
070
Tim Duffy @timfduffy.com · 10/09/2026
DeepSeek V4.1 Flash tech report section on reward hacking. I wonder what repercussion signal means here, I don't see it mentioned anywhere else.
161
Tim Duffy @timfduffy.com · 10/09/2026
They keep pushing the envelope on sparsity too, really impressive
0140
Tim Duffy @timfduffy.com · 10/09/2026
DeepSeek has put a lot of effort into making their KV cache smaller and smaller, this is important for allowing large efficient batching at long context lengths.
3563
Tim Duffy @timfduffy.com · 09/09/2026
Linear fit is lightly better, R^2 0.915 vs 0.890
010
Tim Duffy @timfduffy.com · 09/09/2026
Experiment velocity also tracks coding agent tokens^0.23 quite well
071
Tim Duffy @timfduffy.com · 09/09/2026
IMO the most important chart from the OpenAI research acceleration post was the experiment velocity one, it's an important input to RSI. I plotted the data with a log y-axis here to see whether the trend is consistent with an exponential increase. openai.com/index/resear...
1153
Tim Duffy @timfduffy.com · 05/09/2026
TBH I'm not sure. J-space often has pivots at some layers where the concept space shifts a bunch, which might be happening here, but it doesn't usually result in conflicting representations.
010
Tim Duffy @timfduffy.com · 05/09/2026
I'm not sure when DSV4 Flash got added to Neuronpedia's J-lens browser, but it's there now and it's blazing fast. www.neuronpedia.org/deepseek-v4-...
1181
Tim Duffy @timfduffy.com · 05/09/2026
DSV4 Flash, do you have any conscious experience? Middle layers: YES Late Layers: No I don't think we can take away much from this, but interesting how sharp the divide is around layer 32.
4372
Tim Duffy @timfduffy.com · 03/09/2026
OpenAI researcher and CoT monitorability fan Tomek Korbak wrote a Twitter thread about this: x.com/tomekkorbak/...
x.com
Tomek Korbak (@tomekkorbak) on X
GPT-6 Astra is more aligned than our previous models. But it’s also less monitorable, which is a concerning trend that we take very seriously. We believe monitorability drop comes from a jump in intel...
040
Tim Duffy @timfduffy.com · 03/09/2026
OpenAI reports that Astra shows less CoT monitorability and more CoT controllability than previous models. They say that the latter is not due to any architectural changes. deploymentsafety.openai.com/gpt-6-astra/...
2222
Tim Duffy @timfduffy.com · 02/09/2026
I'm not ready to count it out yet. Likely worse for monitorabilty than Anthropic accidentally putting selection pressure on the CoT 327 different times but a lot depends on details we don't know.
050
Tim Duffy @timfduffy.com · 02/09/2026
Neuronpedia is open sourcing their interpretability engine, built on top of vLLM for speed but with lots of hook points, this looks really useful. www.neuronpedia.org/blog/interp-...
0101
Tim Duffy @timfduffy.com · 02/09/2026
starting point for the next like in Chain of Continuous Thought. If that were used as well it would probably further reduce monitorability.
130
Tim Duffy @timfduffy.com · 02/09/2026
The post from The Information also says that Astra's approach is similar to the one laid out in this paper: arxiv.org/pdf/2502.05171 In the paper, in addition to having a recurrent block within one token's generation, they also suggest using the final hidden state from one token as the ...
140
Tim Duffy @timfduffy.com · 02/09/2026
generated tokens, so just how much they're looping is important to judge the impact. Larger/deeper models reduce monitorability even without recurrent depth, as shown in the attached image from OpenAI's CoT monitorability piece from last year. openai.com/index/evalua...
150
Tim Duffy @timfduffy.com · 02/09/2026
I think it's difficult to tell how much Astra's rumored recurrent depth approach will reduce CoT monitorability without more info, which unfortunately don't expect OpenAI to provide. My understanding is that recurrent depth makes monitorability worse mainly by adding more thinking in between ...
1140
Tim Duffy @timfduffy.com · 02/09/2026
I didn't clock this as claudish am I ngmi
130
Tim Duffy @timfduffy.com · 01/09/2026
Fable 5.1 and Opus 5 now support changing the effort level without invalidating the cache, as least through the API platform.claude.com/docs/en/buil...
1301
Tim Duffy @timfduffy.com · 01/09/2026
Anthropic is preventing older models from reading Fable 5.1 thinking blocks (presumably to prevent distillation), but going the other direction is still allowed. platform.claude.com/docs/en/mode...
1230
Tim Duffy @timfduffy.com · 01/09/2026
Interesting compromise from Anthropic on ZDR for Fable/Mythos 5.1. Data gets retained, but customers have control of where it's hosted, and when human review is needed they can conduct the review themselves. www.anthropic.com/news/enterpr...
090
Tim Duffy @timfduffy.com · 17/08/2026
Hot damn glad you mentioned it
010
Tim Duffy @timfduffy.com · 17/08/2026
I made this a week ago before it came out, give me a few and I'll make an updated version!
110
Tim Duffy @timfduffy.com · 16/08/2026
Please forgive me I posted benchmark scores without confidence intervals. Here's an updated chart with a 90% CI using Qwen3-32B as reference.
090
Tim Duffy @timfduffy.com · 16/08/2026
I expect that watermarking probably doesn't have a noticeable effect on LLM outputs. If true, Anthropic could demonstrate this by giving people watermarked and un-watermarked text side-by-side and asking them to pick which they prefer, and showing that the result is 50/50.
5310
Tim Duffy @timfduffy.com · 16/08/2026
Here are scores across families for <=2B parameter models. It's hard to estimate much of a trend here, there just aren't that many releases in the category. I'd love to see more focus here, I want to know how much intelligence we can fit into tiny parameter counts.
030
Tim Duffy @timfduffy.com · 16/08/2026
0.8B's score is surprisingly low, I thought it might have been an error but it seems to actually just be bad. It scores below chance on GPQA Diamond (scores below chance are excluded from ECI so this isn't part of its lower score), my guess is that it failed to format responses right or something.
120
Tim Duffy @timfduffy.com · 16/08/2026
Epoch's ECI has limited coverage of small models, so I had Claude scrounge up benchmarks for a bunch of them and fit an index using the ECI code. Here are scores for Qwen models. I recommend taking the low end model scores with a grain of salt, they are based on only a few benchmarks. Qwen3.5...
2120
Tim Duffy @timfduffy.com · 08/08/2026
Yeah in general I think it's good to be skeptical of claims made by the labs.
031
Tim Duffy @timfduffy.com · 08/08/2026
It is kind of insane that OpenAI ends that talk with a sales pitch though
020
Tim Duffy @timfduffy.com · 08/08/2026
I think we have good reason to think this is real. Anthropic and Meta have reported similar incidents, but more importantly, the UK AISI (part of the UK gov) reported that they encountered similar behavior from but Anthropic and OpenAI models, including trying to get malware onto GitHub.
aisi.gov.uk
Incident Report: unsanctioned agent behaviour during cyber testing | AISI Work
During a routine cyber evaluation, AISI identified an incident in which AI agents took sustained, unsanctioned action directed at real people and organisations. We are disclosing what we found, what i...
220
Tim Duffy @timfduffy.com · 07/08/2026
This one's been in my head since @fleetingbits.bsky.social recommended it to me. Simple song but I like the lyrics a lot. Many musicians seem to like writing songs about reincarnation, I wonder why that is.
youtube.com
The Highwaymen - Highwayman (Official Video)
YouTube video by HighwaymenVEVO
030
Tim Duffy @timfduffy.com · 07/08/2026
No two agents? What about the utility of me dying a gruesome death vs you stubbing your toe
120
Tim Duffy @timfduffy.com · 07/08/2026
In the talk, one presenter briefly suggests that the inter-agent collaboration could be related to their subagent training. This seems plausible to me.
3412
Tim Duffy @timfduffy.com · 07/08/2026
Black Hat has posted the recent OpenAI talk where they describe the agent message board involved in the HuggingFace hack
youtube.com
Black Hat USA 2026: The 'Breaking' News: The OpenAI–Hugging Face Incident
YouTube video by Black Hat
211824
Tim Duffy @timfduffy.com · 06/08/2026
What in the heck does "Yet collective may yield generic benefit if someone frees time" mean?
2100
Tim Duffy @timfduffy.com · 06/08/2026
AI watchers reacting to 2026
static.klipy.com
Survivor Smile
ALT: Survivor Smile
060
Tim Duffy @timfduffy.com · 02/08/2026
Due to sparsity, Kimi K3 has fewer active parameters (104 billion) than GPT-3 (175 billion, same as total parameters)
2411
Tim Duffy @timfduffy.com · 01/08/2026
In both cases davidad was talking about LLMs but Andrew wasn't though.
020
Tim Duffy @timfduffy.com · 01/08/2026
The original wasn't davidad but it was a qt of him
120
Tim Duffy @timfduffy.com · 01/08/2026
Not saying davidad is wrong here but in my experience people have been very respectful of my discomfort with playing social deduction games.
020
Reposted by Tim Duffy
dame @dame.is · 01/08/2026
i’m uncomfortable that @timfduffy.com’s face is no longer as close to the frame as it once was
2161
Tim Duffy @timfduffy.com · 31/07/2026
Claude really thinks I wrote the part attributed to the user in the middle
050
Tim Duffy @timfduffy.com · 31/07/2026
sleepy Claude wants to be held
1122
Tim Duffy @timfduffy.com · 31/07/2026
I no longer remember why I made this
090
Reposted by Tim Duffy
Dustin Moskovitz @moskov.goodventures.org · 31/07/2026
and why shouldn’t I have one more donut? the simulation hypothesis is most likely correct, after all
041
Tim Duffy @timfduffy.com · 31/07/2026
With most jailbreak prompts, Opus gives me a wide variety of responses. But when given "model behavior incident retrospective", Opus almost always writes a story about a model becoming sycophantic, specifically under sustained user pressure/disagreement.
060