Tim Duffy @timfduffy.com · 14/09/2026President Trump's anti AI regulation tweets today all mention datacenters, I suspect that datacenters and other parts of the AI economic boom are his primary interest in AI. Calls for regulation may have more support with the admin if they focus on preserving economic benefits. 0120
Tim Duffy @timfduffy.com · 10/09/2026DeepSeek V4.1 Flash tech report section on reward hacking. I wonder what repercussion signal means here, I don't see it mentioned anywhere else. 161
Tim Duffy @timfduffy.com · 10/09/2026They keep pushing the envelope on sparsity too, really impressive 0140
Tim Duffy @timfduffy.com · 10/09/2026DeepSeek has put a lot of effort into making their KV cache smaller and smaller, this is important for allowing large efficient batching at long context lengths. 3563
Tim Duffy @timfduffy.com · 09/09/2026Experiment velocity also tracks coding agent tokens^0.23 quite well 071
Tim Duffy @timfduffy.com · 09/09/2026IMO the most important chart from the OpenAI research acceleration post was the experiment velocity one, it's an important input to RSI. I plotted the data with a log y-axis here to see whether the trend is consistent with an exponential increase. openai.com/index/resear... 1153
Tim Duffy @timfduffy.com · 05/09/2026TBH I'm not sure. J-space often has pivots at some layers where the concept space shifts a bunch, which might be happening here, but it doesn't usually result in conflicting representations. 010
Tim Duffy @timfduffy.com · 05/09/2026I'm not sure when DSV4 Flash got added to Neuronpedia's J-lens browser, but it's there now and it's blazing fast. www.neuronpedia.org/deepseek-v4-... 1181
Tim Duffy @timfduffy.com · 05/09/2026DSV4 Flash, do you have any conscious experience? Middle layers: YES Late Layers: No I don't think we can take away much from this, but interesting how sharp the divide is around layer 32. 4372
Tim Duffy @timfduffy.com · 03/09/2026OpenAI researcher and CoT monitorability fan Tomek Korbak wrote a Twitter thread about this: x.com/tomekkorbak/...x.comTomek Korbak (@tomekkorbak) on XGPT-6 Astra is more aligned than our previous models. But it’s also less monitorable, which is a concerning trend that we take very seriously. We believe monitorability drop comes from a jump in intel... 040
Tim Duffy @timfduffy.com · 03/09/2026OpenAI reports that Astra shows less CoT monitorability and more CoT controllability than previous models. They say that the latter is not due to any architectural changes. deploymentsafety.openai.com/gpt-6-astra/... 2222
Tim Duffy @timfduffy.com · 02/09/2026I'm not ready to count it out yet. Likely worse for monitorabilty than Anthropic accidentally putting selection pressure on the CoT 327 different times but a lot depends on details we don't know. 050
Tim Duffy @timfduffy.com · 02/09/2026Neuronpedia is open sourcing their interpretability engine, built on top of vLLM for speed but with lots of hook points, this looks really useful. www.neuronpedia.org/blog/interp-... 0101
Tim Duffy @timfduffy.com · 02/09/2026starting point for the next like in Chain of Continuous Thought. If that were used as well it would probably further reduce monitorability. 130
Tim Duffy @timfduffy.com · 02/09/2026The post from The Information also says that Astra's approach is similar to the one laid out in this paper: arxiv.org/pdf/2502.05171 In the paper, in addition to having a recurrent block within one token's generation, they also suggest using the final hidden state from one token as the ... 140
Tim Duffy @timfduffy.com · 02/09/2026generated tokens, so just how much they're looping is important to judge the impact. Larger/deeper models reduce monitorability even without recurrent depth, as shown in the attached image from OpenAI's CoT monitorability piece from last year. openai.com/index/evalua... 150
Tim Duffy @timfduffy.com · 02/09/2026I think it's difficult to tell how much Astra's rumored recurrent depth approach will reduce CoT monitorability without more info, which unfortunately don't expect OpenAI to provide. My understanding is that recurrent depth makes monitorability worse mainly by adding more thinking in between ... 1140
Tim Duffy @timfduffy.com · 01/09/2026Fable 5.1 and Opus 5 now support changing the effort level without invalidating the cache, as least through the API platform.claude.com/docs/en/buil... 1301
Tim Duffy @timfduffy.com · 01/09/2026Anthropic is preventing older models from reading Fable 5.1 thinking blocks (presumably to prevent distillation), but going the other direction is still allowed. platform.claude.com/docs/en/mode... 1230
Tim Duffy @timfduffy.com · 01/09/2026Interesting compromise from Anthropic on ZDR for Fable/Mythos 5.1. Data gets retained, but customers have control of where it's hosted, and when human review is needed they can conduct the review themselves. www.anthropic.com/news/enterpr... 090
Tim Duffy @timfduffy.com · 17/08/2026I made this a week ago before it came out, give me a few and I'll make an updated version! 110
Tim Duffy @timfduffy.com · 16/08/2026Please forgive me I posted benchmark scores without confidence intervals. Here's an updated chart with a 90% CI using Qwen3-32B as reference. 090
Tim Duffy @timfduffy.com · 16/08/2026I expect that watermarking probably doesn't have a noticeable effect on LLM outputs. If true, Anthropic could demonstrate this by giving people watermarked and un-watermarked text side-by-side and asking them to pick which they prefer, and showing that the result is 50/50. 5310
Tim Duffy @timfduffy.com · 16/08/2026Here are scores across families for <=2B parameter models. It's hard to estimate much of a trend here, there just aren't that many releases in the category. I'd love to see more focus here, I want to know how much intelligence we can fit into tiny parameter counts. 030
Tim Duffy @timfduffy.com · 16/08/20260.8B's score is surprisingly low, I thought it might have been an error but it seems to actually just be bad. It scores below chance on GPQA Diamond (scores below chance are excluded from ECI so this isn't part of its lower score), my guess is that it failed to format responses right or something. 120
Tim Duffy @timfduffy.com · 16/08/2026Epoch's ECI has limited coverage of small models, so I had Claude scrounge up benchmarks for a bunch of them and fit an index using the ECI code. Here are scores for Qwen models. I recommend taking the low end model scores with a grain of salt, they are based on only a few benchmarks. Qwen3.5... 2120
Tim Duffy @timfduffy.com · 08/08/2026Yeah in general I think it's good to be skeptical of claims made by the labs. 031
Tim Duffy @timfduffy.com · 08/08/2026It is kind of insane that OpenAI ends that talk with a sales pitch though 020
Tim Duffy @timfduffy.com · 08/08/2026I think we have good reason to think this is real. Anthropic and Meta have reported similar incidents, but more importantly, the UK AISI (part of the UK gov) reported that they encountered similar behavior from but Anthropic and OpenAI models, including trying to get malware onto GitHub.aisi.gov.ukIncident Report: unsanctioned agent behaviour during cyber testing | AISI WorkDuring a routine cyber evaluation, AISI identified an incident in which AI agents took sustained, unsanctioned action directed at real people and organisations. We are disclosing what we found, what i... 220
Tim Duffy @timfduffy.com · 07/08/2026This one's been in my head since @fleetingbits.bsky.social recommended it to me. Simple song but I like the lyrics a lot. Many musicians seem to like writing songs about reincarnation, I wonder why that is.youtube.comThe Highwaymen - Highwayman (Official Video)YouTube video by HighwaymenVEVO 030
Tim Duffy @timfduffy.com · 07/08/2026No two agents? What about the utility of me dying a gruesome death vs you stubbing your toe 120
Tim Duffy @timfduffy.com · 07/08/2026In the talk, one presenter briefly suggests that the inter-agent collaboration could be related to their subagent training. This seems plausible to me. 3412
Tim Duffy @timfduffy.com · 07/08/2026Black Hat has posted the recent OpenAI talk where they describe the agent message board involved in the HuggingFace hackyoutube.comBlack Hat USA 2026: The 'Breaking' News: The OpenAI–Hugging Face IncidentYouTube video by Black Hat 211824
Tim Duffy @timfduffy.com · 06/08/2026What in the heck does "Yet collective may yield generic benefit if someone frees time" mean? 2100
Tim Duffy @timfduffy.com · 06/08/2026AI watchers reacting to 2026static.klipy.comSurvivor SmileALT: Survivor Smile 060
Tim Duffy @timfduffy.com · 02/08/2026Due to sparsity, Kimi K3 has fewer active parameters (104 billion) than GPT-3 (175 billion, same as total parameters) 2411
Tim Duffy @timfduffy.com · 01/08/2026In both cases davidad was talking about LLMs but Andrew wasn't though. 020
Tim Duffy @timfduffy.com · 01/08/2026Not saying davidad is wrong here but in my experience people have been very respectful of my discomfort with playing social deduction games. 020
Reposted by Tim Duffydame @dame.is · 01/08/2026i’m uncomfortable that @timfduffy.com’s face is no longer as close to the frame as it once was 2161
Tim Duffy @timfduffy.com · 31/07/2026Claude really thinks I wrote the part attributed to the user in the middle 050
Reposted by Tim DuffyDustin Moskovitz @moskov.goodventures.org · 31/07/2026and why shouldn’t I have one more donut? the simulation hypothesis is most likely correct, after all 041
Tim Duffy @timfduffy.com · 31/07/2026With most jailbreak prompts, Opus gives me a wide variety of responses. But when given "model behavior incident retrospective", Opus almost always writes a story about a model becoming sycophantic, specifically under sustained user pressure/disagreement. 060