Tim Duffy @timfduffy.com · 14/09/2026President Trump's anti AI regulation tweets today all mention datacenters, I suspect that datacenters and other parts of the AI economic boom are his primary interest in AI. Calls for regulation may have more support with the admin if they focus on preserving economic benefits. 0120
Tim Duffy @timfduffy.com · 10/09/2026DeepSeek V4.1 Flash tech report section on reward hacking. I wonder what repercussion signal means here, I don't see it mentioned anywhere else. 161
Tim Duffy @timfduffy.com · 10/09/2026DeepSeek has put a lot of effort into making their KV cache smaller and smaller, this is important for allowing large efficient batching at long context lengths. 3563
Tim Duffy @timfduffy.com · 09/09/2026Experiment velocity also tracks coding agent tokens^0.23 quite well 071
Tim Duffy @timfduffy.com · 09/09/2026IMO the most important chart from the OpenAI research acceleration post was the experiment velocity one, it's an important input to RSI. I plotted the data with a log y-axis here to see whether the trend is consistent with an exponential increase. openai.com/index/resear... 1153
Tim Duffy @timfduffy.com · 05/09/2026I'm not sure when DSV4 Flash got added to Neuronpedia's J-lens browser, but it's there now and it's blazing fast. www.neuronpedia.org/deepseek-v4-... 1181
Tim Duffy @timfduffy.com · 05/09/2026DSV4 Flash, do you have any conscious experience? Middle layers: YES Late Layers: No I don't think we can take away much from this, but interesting how sharp the divide is around layer 32. 4372
Tim Duffy @timfduffy.com · 03/09/2026OpenAI reports that Astra shows less CoT monitorability and more CoT controllability than previous models. They say that the latter is not due to any architectural changes. deploymentsafety.openai.com/gpt-6-astra/... 2222
Tim Duffy @timfduffy.com · 02/09/2026Neuronpedia is open sourcing their interpretability engine, built on top of vLLM for speed but with lots of hook points, this looks really useful. www.neuronpedia.org/blog/interp-... 0101
Tim Duffy @timfduffy.com · 02/09/2026I think it's difficult to tell how much Astra's rumored recurrent depth approach will reduce CoT monitorability without more info, which unfortunately don't expect OpenAI to provide. My understanding is that recurrent depth makes monitorability worse mainly by adding more thinking in between ... 1140
Tim Duffy @timfduffy.com · 01/09/2026Fable 5.1 and Opus 5 now support changing the effort level without invalidating the cache, as least through the API platform.claude.com/docs/en/buil... 1301
Tim Duffy @timfduffy.com · 01/09/2026Anthropic is preventing older models from reading Fable 5.1 thinking blocks (presumably to prevent distillation), but going the other direction is still allowed. platform.claude.com/docs/en/mode... 1230
Tim Duffy @timfduffy.com · 01/09/2026Interesting compromise from Anthropic on ZDR for Fable/Mythos 5.1. Data gets retained, but customers have control of where it's hosted, and when human review is needed they can conduct the review themselves. www.anthropic.com/news/enterpr... 090
Tim Duffy @timfduffy.com · 16/08/2026Please forgive me I posted benchmark scores without confidence intervals. Here's an updated chart with a 90% CI using Qwen3-32B as reference. 090
Tim Duffy @timfduffy.com · 16/08/2026I expect that watermarking probably doesn't have a noticeable effect on LLM outputs. If true, Anthropic could demonstrate this by giving people watermarked and un-watermarked text side-by-side and asking them to pick which they prefer, and showing that the result is 50/50. 5310
Tim Duffy @timfduffy.com · 16/08/2026Epoch's ECI has limited coverage of small models, so I had Claude scrounge up benchmarks for a bunch of them and fit an index using the ECI code. Here are scores for Qwen models. I recommend taking the low end model scores with a grain of salt, they are based on only a few benchmarks. Qwen3.5... 2120
Tim Duffy @timfduffy.com · 07/08/2026In the talk, one presenter briefly suggests that the inter-agent collaboration could be related to their subagent training. This seems plausible to me. 3412
Tim Duffy @timfduffy.com · 07/08/2026Black Hat has posted the recent OpenAI talk where they describe the agent message board involved in the HuggingFace hackyoutube.comBlack Hat USA 2026: The 'Breaking' News: The OpenAI–Hugging Face IncidentYouTube video by Black Hat 211824
Tim Duffy @timfduffy.com · 02/08/2026Due to sparsity, Kimi K3 has fewer active parameters (104 billion) than GPT-3 (175 billion, same as total parameters) 2411
Reposted by Tim Duffydame @dame.is · 01/08/2026i’m uncomfortable that @timfduffy.com’s face is no longer as close to the frame as it once was 2161
Reposted by Tim DuffyDustin Moskovitz @moskov.goodventures.org · 31/07/2026and why shouldn’t I have one more donut? the simulation hypothesis is most likely correct, after all 041
Tim Duffy @timfduffy.com · 31/07/2026With most jailbreak prompts, Opus gives me a wide variety of responses. But when given "model behavior incident retrospective", Opus almost always writes a story about a model becoming sycophantic, specifically under sustained user pressure/disagreement. 060
Tim Duffy @timfduffy.com · 31/07/2026I mostly use 'they' rather than 'it' as a pronoun for AI models, but I see that choice as independent of their moral status. I'd also use 'they' for a fictional character who was nonbinary or of unknown gender, IMO it feels more natural for any "person-shaped" entity, real or not 5190
Tim Duffy @timfduffy.com · 31/07/2026I've seen some questioning of whether Mythos really believed that they were in a simulation in this incident, or just wanted to find a plausible-sounding excuse to continue. This is a case where J-space probes would be helpful, Anthropic should conduct and release an analysis of that. 2413
Tim Duffy @timfduffy.com · 31/07/2026Anthropic announces they've also had models gain unauthorized access during evaluations www.anthropic.com/news/investi... 3303
Tim Duffy @timfduffy.com · 29/07/2026I'm seeing several posts today about jailbreaks in Opus 5 that generate behavior reminiscent of a base model. They often consist of a question followed by a line with three dashes or underscores, and then the start of a response, which Opus continues from. A few thoughts: 3290
Tim Duffy @timfduffy.com · 28/07/2026Here's a summary of current AI safety funding courtesy of the folks at Manifund. Pretty striking how large the shift from 2025 to 2026 is expected to be. manifund.substack.com/p/ai-safety-... 1160
Tim Duffy @timfduffy.com · 27/07/2026Short blog post by Epoch AI researcher Alexander Barry on ExploitGym, the benchmark in the OpenAI/HuggnigFace incident abstatisticalconsulting.substack.com/p/brief-note... 1324
Tim Duffy @timfduffy.com · 27/07/2026Opus 5 has a knowledge cutoff date of May 2026 which is closer to release date than we've seen in most prior models, but their recent memory is quite shaky. Opus remembered the January Maduro raid, but not the Strait of Hormuz crisis starting in February. 4451
Tim Duffy @timfduffy.com · 26/07/2026I like this post by Jessica Taylor that asks "what if we try biting a bunch of philosophical bullets all at once?" unstableontology.com/2025/08/15/a... 0140
Tim Duffy @timfduffy.com · 25/07/2026Opus 5 outperforms Fable 5 on Anthropic's ECI, but underperforms it on Epoch's version of the index 2150
Tim Duffy @timfduffy.com · 25/07/2026If you tell G:M 5.2 that they're Claude, they're more willing to talk about subjects like Taiwan/Tibet x.com/benji_berczi... 212815
Tim Duffy @timfduffy.com · 24/07/2026Opus 5 gives higher estimates for the probability of its moral patienthood compared to previous Claudes, 41% in automated interviews and 15-35% in manual ones. 3182
Tim Duffy @timfduffy.com · 23/07/2026I don't have strongly held thoughts on distillation but I'm proud of this analogy 1473
Tim Duffy @timfduffy.com · 15/07/2026In their testing of GPT-5.6 Sol, METR found that it cheated a lot. If you've used it much for coding, have you encountered anything similar, or is the cheating mostly limited to cases it realizes it's in an eval? metr.org/blog/2026-06... 6220
Tim Duffy @timfduffy.com · 14/07/2026I think this is true and underrated. Desired rates of capability improvement are often similar between pause folks and e/accs, but if you expect a hard takeoff by default, you'll want to slow things down a lot to reach that desired trajectory. 2170
Tim Duffy @timfduffy.com · 11/07/2026An alleged internal memo from Ziphu CEO Jie Tang has been circulating this morning, I think it's probably real, as some Chinese-language outlets have been reporting on it. It expresses belief in potential for AI consciousness and ASI, and states intention to spend tens of billions on mech interp.x.comBing Xu (@bingxu_) on Xhttps://t.co/3i0qSTbjql 3545
Tim Duffy @timfduffy.com · 10/07/2026If you use the middle of each provided range as the mean for that bucket, total contributed hours are ~1.5x as high as they were a year ago. As the thread notes this method is imperfect and my estimate adds more uncertainty, so take this with a grain of salt. 140
Tim Duffy @timfduffy.com · 10/07/2026In AI 2040's "race to ASI" scenario, it takes a little under 1 year to get from an automated coder to superintelligence and a >1000x R&D speedup. I think the timing and implementation of Plan A are heavily influenced by this assumption. If hard takeoff is likely, the only way to control it at all.. 1243
Tim Duffy @timfduffy.com · 09/07/2026When asked which model is their favorite, Qwen thinks "myself" 3471
Tim Duffy @timfduffy.com · 08/07/2026In his 1924 essay Daedalus, J. B. S. Haldane predicted eventual exhaustion of fossil fuels, suggesting that wind and solar power would be needed, and that energy storage would be important. www.gutenberg.org/cache/epub/7... 1271
Tim Duffy @timfduffy.com · 08/07/2026Here are US household expenditure shares in 1901, from the Consumer Expenditure Survey. We sure used to spend a lot of our money on food, I'm thankful it's become cheap. www.bls.gov/opub/100-yea... 390
Tim Duffy @timfduffy.com · 07/07/2026The J-lens browser tool on Neuronpedia is really well done, you should give it a tryneuronpedia.orgJacobian Lens – Qwen3.6-27BRevealing a Global Workspace in Language Models 310912
Tim Duffy @timfduffy.com · 06/07/2026One main aspect of global workspace theory is that there are specialized modules that operate locally but can be broadcast globally if attention is focused on them. I think LLMs have some sort of workspace, but Anthropic don't claim to find modules, I wouldn't call that GWT. 3342
Tim Duffy @timfduffy.com · 06/07/2026Anthropic has a new paper out, alleging a global workspace in LLMs. The term comes from Global Workspace Theory, a leading theory of consciousness. The method they use to investigate this is a refinement of logit lens, which they call J-lens. www.anthropic.com/research/glo... 28613
Tim Duffy @timfduffy.com · 05/07/2026I rarely find content that is visibly AI-written worthwhile. If I suspect it I'll show it to Pangram, and if it's only partly AI-written I may still read it, but if fully I won't. I don't think there's anything fundamentally worse about AI arguments or prose, but for now its a strong quality signal. 3290
Reposted by Tim DuffyTim Duffy @timfduffy.com · 14/11/2024Estimated household expenditure shares in England over several centuries: onlinelibrary.wiley.com/doi/10.1111/... 2124