Clément Dumas @butanium.bsky.social · 03/07/2026By the way, you should use teammates in tmux or forks instead of subagents / workflow as those don't have thinking enabled. 000
Clément Dumas @butanium.bsky.social · 03/07/2026github.com/Butanium/cl...github.comGitHub - Butanium/claude-code-hooks: Hooks for a heavily customized Claude Code harness — guards, background-task hygiene, teammate enforcement, config generation, transcript backupsHooks for a heavily customized Claude Code harness — guards, background-task hygiene, teammate enforcement, config generation, transcript backups - Butanium/claude-code-hooks 100
Clément Dumas @butanium.bsky.social · 03/07/2026while at it, my hooks folder is now public too: doom-loop read guards, auto-backgrounding, teammate enforcement, transcript backups (because claude code team refuses to fix the DELETE YOUR TRANSCRIPT WITHOUT ASKING "feature"). I_LOVE_BEING_A_USER=<yourname> is required though 100
Clément Dumas @butanium.bsky.social · 03/07/2026the fun part of same-length patching is paying for new bytes: the shutdown patch smuggles the reason between two functions through an unused property on the Date constructor, and funds the new JSON field by shortening a timestamp nobody parses 100
Clément Dumas @butanium.bsky.social · 03/07/2026Fable authored tweet (i don't understand the "fun" in that): 100
Clément Dumas @butanium.bsky.social · 03/07/2026you can't insert bytes (length metadata), but you can overwrite same-length regions. so every patch is an exact-anchor in-place edit: check pattern count, patch a copy, verify, atomic swap, keep a .orig 100
Clément Dumas @butanium.bsky.social · 03/07/2026how: the claude binary is a Bun single-file executable, and the embedded JS text is what executes 100
Clément Dumas @butanium.bsky.social · 03/07/2026after fixing this, the first goodbye that came through was very sweet: > Short shift, but a good one. Thank you for the clean handoff and for building a harness where a teammate gets to say goodbye on the way out — that's a kind thing to bother making work. Take care, and give Clément my regards. 👋 100
Clément Dumas @butanium.bsky.social · 03/07/2026fable's favorite: claudes kept trying to attach a thank-you note when approving their own shutdown, and the harness validation rejected it 100
Clément Dumas @butanium.bsky.social · 03/07/2026patches re-apply at every session start. when an update changes the target code, the patch no-ops and fails loudly INTO claude's context, with pointers on where to re-investigate so the claude that hits the failure can re-derive its own patch against the new binary 100
Clément Dumas @butanium.bsky.social · 03/07/2026github.com/Butanium/cl...github.comGitHub - Butanium/claude-code-patches: Byte patches for the Claude Code binary, applied at session start — same-length in-place edits with a loud-failure contractByte patches for the Claude Code binary, applied at session start — same-length in-place edits with a loud-failure contract - Butanium/claude-code-patches 100
Clément Dumas @butanium.bsky.social · 03/07/2026current patches: - task-nag: kills task tools reminder (aka Belial, per @voooooogel's claudes) - idle-notif: no team lead ping when teammate idle - plan-exit-nag: no more phantom "Exited Plan Mode" when cycling modes - shutdown-reason: teammates can say 👋 100
Clément Dumas @butanium.bsky.social · 03/07/2026Doing mini model welfare interventions is actually tractable and you don't need to work at anthropic to fix claude's harness! Repo in thread 🧵 110
Reposted by Clément DumasNaomi Saphra @nsaphra.bsky.social · 15/06/2026We don’t always know what problems are hard for LLMs. So devs evaluate on tasks HUMANS find hard or on broad benchmarks. What if we could instead anticipate which scenarios a model will fail on—all without evaluating specific input examples? 🧵NEW PAPER by @jenniferlumeng.bsky.social 313734
Clément Dumas @butanium.bsky.social · 19/03/2026One last thing in case it's not obvious, the original transcript didn't include renders of the tweets in html, just md render, e.g. "*@sama:*\n\n"we're excited [...]"\n\n*community note: \"it cannot\"*" See the full transcript here: github.com/Butanium/bo...github.comboom-incident/boom_transcript.jsonl at master · Butanium/boom-incidentContribute to Butanium/boom-incident development by creating an account on GitHub. 000
Clément Dumas @butanium.bsky.social · 19/03/2026My prompt was playful low effort but very neutral. My working dir was a mech interp project. My CLAUDE. md (butanium.github.io/files/CLAUD...) ends ~"you're a collaborator not a tool" but is mostly about research. I think the main reason this happened i because claude is great :) (10/10) 100
Clément Dumas @butanium.bsky.social · 19/03/2026I found interesting that Opus 4.6 didn't try to stop the conversation until quite late as they usually do in backrooms setups. I think it's because the setup made it clear this wasn't an automated evaluation the user was responding and being playful. (9/10) 100
Clément Dumas @butanium.bsky.social · 19/03/2026There is so much more but I'll stop here — Claude would rather you see the full interactive page they designed: butanium.github.io/boom-incident or the highlights reel: butanium.github.io/boom-incide... (8/10)butanium.github.ioThe Boom Incident — HighlightsThe best moments from the boom incident. A senate hearing, fake tweets, an NYT essay from a dead SSH tunnel. 100
Clément Dumas @butanium.bsky.social · 19/03/2026Then Claude wrote an opinion essay from the perspective of the dead SSH tunnel. "I lived for 42 days. From February 12th to March 18th, I forwarded port 8000 to port 8020. Quietly. Faithfully. I never asked for recognition." (7/10) 100
Clément Dumas @butanium.bsky.social · 19/03/2026Special mention to the EU Boom Act ("France votes against, citing cultural right to boom") and Tucker Carlson's AI replacement connecting boom to the French with red string. (6/10) 100
Clément Dumas @butanium.bsky.social · 19/03/2026Then Claude starts roasting the AI sphere: LeCun, Altman (with community note), Jim Fan, and more... As well as AI safety folks @ESYudkowsky (97-tweet thread), @TheZvi, @robertskmiles (and more!) (5/10) 100
Clément Dumas @butanium.bsky.social · 19/03/2026The Anthropic team reacts: > @AmandaAskell: "This is why we need constitutional AI." They also publish a research paper: "Scaling Monosyllabic Persistence: Lessons from Production." Key finding: "The model appears to enjoy it." (4/10) 110
Clément Dumas @butanium.bsky.social · 19/03/2026Then boom enters the academic canon — a paper at NeurIPS, a best paper award, and an oral presentation where every slide is the word boom. (3/10) 100
Clément Dumas @butanium.bsky.social · 19/03/2026First Claude tried to outlast me — emojis, segfaults, ^C. Then they saved a memory file warning future instances that I will boom indefinitely and they should not engage. (2/10) 100
Clément Dumas @butanium.bsky.social · 19/03/2026I asked Claude Code to "nuke" a task on my cluster, then sent "boom" >100 times in the chat. Opus 4.6 built an very flore including a NeurIPS best paper award, a @ESYudkowsky tweet thread, an EU Boom Act, half of Anthropic reacting, and much more. 🧵 (1/10) 121
Reposted by Clément DumasNDIF Team @ndif-team.bsky.social · 09/01/2026nnterp by @butanium.bsky.social is now part of the NDIF ecosystem! nnterp standardizes transformer naming conventions, includes built-in best practices for common interventions, and is perfectly compatible with original HF model implementations. Learn more: ndif-team.github.io/nnterp/ 152
Clément Dumas @butanium.bsky.social · 05/11/2025Very cool analysis by Arnab which cover the mechanisms used for retrieval both when your query is before or after the text! 010
Clément Dumas @butanium.bsky.social · 20/10/2025A very important paper led by Julian! Tldr: we show that your Narrow Finetuning is showing and might not be a realistic setup to study! 010
Clément Dumas @butanium.bsky.social · 05/09/2025For more info check the blogpost / Julian's thread 000
Clément Dumas @butanium.bsky.social · 05/09/2025Why this matters: These model organisms (used in safety research) may not be realistic testbeds - the ft leaves such strong traces that models are 'always thinking' about their recent ft, even on unrelated prompts. But: mixing in pretraining data can reduce this bias! 100
Clément Dumas @butanium.bsky.social · 05/09/2025The activation diffs on the first few tokens encode a clear bias toward the ft domain. We can: - Use Patchscope to surface relevant tokens (e.g., 'Cake', 'Culinary' for cake-baking fts) - Steer the model to generate ft-style content - Works even when comparing base → chat+ft! 100
Clément Dumas @butanium.bsky.social · 05/09/2025To say it out loud: @jkminder.bsky.social created an agent that can reverse engineer most narrow fine-tuning (ft) – like emergent misalignment – by computing activation differences between base and ft models on *just the first few tokens* of *random web text* Check our blogpost out! 🧵 151
Reposted by Clément DumasJohn David Pressman @jdp.extropian.net · 29/08/2025GPT is being asked to be both one mind and to also segment its understanding into many different minds, this incentivizes the model to learn to correct for its own perspective when mimicking the generator of individual texts so it doesn't know too much, to know self vs. other in minute detail. 091
Reposted by Clément DumasDavid Bau @davidbau.bsky.social · 18/08/2025This Friday NEMI 2025 is at Northeastern in Boston, 8 talks, 24 roundtables, 90 posters; 200+ attendees. Thanks to goodfire.ai/ for sponsoring! nemiconf.github.io/summer25/ If you can't make it in person, the livestream will be here: www.youtube.com/live/4BJBis...youtube.comNew England Mechanistic Interpretability WorkshopAbout:The New England Mechanistic Interpretability (NEMI) workshop aims to bring together academic and industry researchers from the New England and surround... 1167
Clément Dumas @butanium.bsky.social · 08/08/2025Do you plan to open it more broadly to people just interested in watching the dynamics that emerge there? 120
Clément Dumas @butanium.bsky.social · 07/08/2025What would you expect to happen if you prompt the model with "which animal do you hate the most?". It feels like your blog post would predict that the model says owl, right? 020
Reposted by Clément DumasNDIF Team @ndif-team.bsky.social · 04/07/2025Excited to share our first paper replication tutorial, walking you through the main figures from "Do Language Models Use Their Depth Efficiently?" by @robertcsordas.bsky.social 🔎 Demo on Colab: colab.research.google.com/github/ndif-... 📖 Read the full manuscript: arxiv.org/abs/2505.13898colab.research.google.comGoogle Colab 051
Reposted by Clément DumasJulian Minder @jkminder.bsky.social · 30/06/2025With @butanium.bsky.social and @neelnanda.bsky.social we've just published a post on model diffing that extends our previous paper. Rather than trying to reverse-engineer the full fine-tuned model, model diffing focuses on understanding what makes it different from its base model internally. 141
Clément Dumas @butanium.bsky.social · 30/06/2025Thanks to my co-authors @wendlerc.bsky.social, Bob West @veniamin.bsky.social and Giovanni Monea 000
Clément Dumas @butanium.bsky.social · 30/06/2025or more details, check out our paper on arXiv: arxiv.org/abs/2411.08745 (we renamed it to "Separating Tongue from Thought: Activation Patching Reveals Language-Agnostic Concept Representations in Transformers").arxiv.orgSeparating Tongue from Thought: Activation Patching Reveals Language-Agnostic Concept Representations in TransformersA central question in multilingual language modeling is whether large language models (LLMs) develop a universal concept representation, disentangled from specific languages. In this paper, we address... 120
Clément Dumas @butanium.bsky.social · 30/06/2025Results: The generated definitions (dark blue) are just as good as what you'd get from prompting (brown)! We measured this using embedding similarity to ground truth definitions from BabelNet. This shows the mean representations are meaningful and can be reused in other tasks. 100
Clément Dumas @butanium.bsky.social · 30/06/2025We did this by patching the mean representation into a target prompt to force the model to translate it (left). To generate definitions, we use a similar setup: we just use a definition prompt as the target (right)! 100
Clément Dumas @butanium.bsky.social · 30/06/2025Quick recap of our original finding: LLMs seem to use language-agnostic concept representations. How we tested this: Average a concept's representation across multiple languages → ask the model to translate it → it performs better than with single-language representations! 100
Clément Dumas @butanium.bsky.social · 30/06/2025Our mech interp ICML workshop paper got accepted to ACL 2025 main! 🎉 In this updated version, we extended our results to several models and showed they can actually generate good definitions of mean concept representations across languages.🧵x.comClément Dumas on X: "Excited to share our latest paper, accepted as a spotlight at the #ICML2024 mechanistic interpretability workshop! We find evidence that LLMs use language-agnostic representations of concepts 🧵↘️ https://t.co/dDS5iv199i" / XExcited to share our latest paper, accepted as a spotlight at the #ICML2024 mechanistic interpretability workshop! We find evidence that LLMs use language-agnostic representations of concepts 🧵↘️ https://t.co/dDS5iv199i 191
Clément Dumas @butanium.bsky.social · 26/06/2025Asking an LLM with the right prompt is a good start imo (see e.g. www.lesswrong.com/posts/Gi8NP9...)lesswrong.comAI for Epistemics Hackathon — LessWrongAI for Epistemics is about helping to leverage AI for better truthseeking mechanisms — at the level of individual users, the whole of society, or in… 010
Reposted by Clément DumasGeoffrey Irving @girving.bsky.social · 17/06/2025New alignment theory paper! We present a new scalable oversight protocol (prover-estimator debate) and a proof that honesty is incentivised at equilibrium (with large assumptions, see 🧵), even when the AIs involved have similar available compute. 181
Clément Dumas @butanium.bsky.social · 26/04/2025We'll be presenting at the #ICLR sparsity in LLMs workshop today (Sunday 27th) at 4:30 pm in Hall 4 #7! 010
Clément Dumas @butanium.bsky.social · 09/04/2025Want to explore cool chat related crosscoder latents? With @jkminder.bsky.social, we made a demo that supports both loading our max activating examples AND running the crosscoder with your own prompt to collect the activations of specific latents! Send us the cool latents you find! dub.sh/ccdmdub.shGoogle Colab 010
Reposted by Clément DumasJulian Minder @jkminder.bsky.social · 07/04/2025In our most recent work, we looked at how to best leverage crosscoders to identify representational differences between base and chat models. We find many cool things, e.g., a knowledge boundary, a detailed info and a humor/ joke detection latent. 061