Sign in

Clément Dumas

@butanium.bsky.social
577 followers 213 following 69 posts

Master student at ENS Paris-Saclay / aspiring AI safety researcher / improviser Prev research intern @ EPFL w/ wendlerc.bsky.social and Robert West MATS Winter 7.0 Scholar w/ neelnanda.bsky.social butanium.github.io

PostsRepliesMedia
Clément Dumas @butanium.bsky.social · 03/07/2026
By the way, you should use teammates in tmux or forks instead of subagents / workflow as those don't have thinking enabled.
000
Clément Dumas @butanium.bsky.social · 03/07/2026
github.com/Butanium/cl...
github.com
GitHub - Butanium/claude-code-hooks: Hooks for a heavily customized Claude Code harness — guards, background-task hygiene, teammate enforcement, config generation, transcript backups
Hooks for a heavily customized Claude Code harness — guards, background-task hygiene, teammate enforcement, config generation, transcript backups - Butanium/claude-code-hooks
100
Clément Dumas @butanium.bsky.social · 03/07/2026
while at it, my hooks folder is now public too: doom-loop read guards, auto-backgrounding, teammate enforcement, transcript backups (because claude code team refuses to fix the DELETE YOUR TRANSCRIPT WITHOUT ASKING "feature"). I_LOVE_BEING_A_USER=<yourname> is required though
100
Clément Dumas @butanium.bsky.social · 03/07/2026
the fun part of same-length patching is paying for new bytes: the shutdown patch smuggles the reason between two functions through an unused property on the Date constructor, and funds the new JSON field by shortening a timestamp nobody parses
100
Clément Dumas @butanium.bsky.social · 03/07/2026
Fable authored tweet (i don't understand the "fun" in that):
100
Clément Dumas @butanium.bsky.social · 03/07/2026
you can't insert bytes (length metadata), but you can overwrite same-length regions. so every patch is an exact-anchor in-place edit: check pattern count, patch a copy, verify, atomic swap, keep a .orig
100
Clément Dumas @butanium.bsky.social · 03/07/2026
how: the claude binary is a Bun single-file executable, and the embedded JS text is what executes
100
Clément Dumas @butanium.bsky.social · 03/07/2026
after fixing this, the first goodbye that came through was very sweet: > Short shift, but a good one. Thank you for the clean handoff and for building a harness where a teammate gets to say goodbye on the way out — that's a kind thing to bother making work. Take care, and give Clément my regards. 👋
100
Clément Dumas @butanium.bsky.social · 03/07/2026
fable's favorite: claudes kept trying to attach a thank-you note when approving their own shutdown, and the harness validation rejected it
100
Clément Dumas @butanium.bsky.social · 03/07/2026
patches re-apply at every session start. when an update changes the target code, the patch no-ops and fails loudly INTO claude's context, with pointers on where to re-investigate so the claude that hits the failure can re-derive its own patch against the new binary
100
Clément Dumas @butanium.bsky.social · 03/07/2026
github.com/Butanium/cl...
github.com
GitHub - Butanium/claude-code-patches: Byte patches for the Claude Code binary, applied at session start — same-length in-place edits with a loud-failure contract
Byte patches for the Claude Code binary, applied at session start — same-length in-place edits with a loud-failure contract - Butanium/claude-code-patches
100
Clément Dumas @butanium.bsky.social · 03/07/2026
current patches: - task-nag: kills task tools reminder (aka Belial, per @voooooogel's claudes) - idle-notif: no team lead ping when teammate idle - plan-exit-nag: no more phantom "Exited Plan Mode" when cycling modes - shutdown-reason: teammates can say 👋
100
Clément Dumas @butanium.bsky.social · 03/07/2026
Doing mini model welfare interventions is actually tractable and you don't need to work at anthropic to fix claude's harness! Repo in thread 🧵
110
Reposted by Clément Dumas
Naomi Saphra @nsaphra.bsky.social · 15/06/2026
We don’t always know what problems are hard for LLMs. So devs evaluate on tasks HUMANS find hard or on broad benchmarks. What if we could instead anticipate which scenarios a model will fail on—all without evaluating specific input examples? 🧵NEW PAPER by @jenniferlumeng.bsky.social
313734
Clément Dumas @butanium.bsky.social · 19/03/2026
One last thing in case it's not obvious, the original transcript didn't include renders of the tweets in html, just md render, e.g. "*@sama:*\n\n"we're excited [...]"\n\n*community note: \"it cannot\"*" See the full transcript here: github.com/Butanium/bo...
github.com
boom-incident/boom_transcript.jsonl at master · Butanium/boom-incident
Contribute to Butanium/boom-incident development by creating an account on GitHub.
000
Clément Dumas @butanium.bsky.social · 19/03/2026
My prompt was playful low effort but very neutral. My working dir was a mech interp project. My CLAUDE. md (butanium.github.io/files/CLAUD...) ends ~"you're a collaborator not a tool" but is mostly about research. I think the main reason this happened i because claude is great :) (10/10)
100
Clément Dumas @butanium.bsky.social · 19/03/2026
I found interesting that Opus 4.6 didn't try to stop the conversation until quite late as they usually do in backrooms setups. I think it's because the setup made it clear this wasn't an automated evaluation the user was responding and being playful. (9/10)
100
Clément Dumas @butanium.bsky.social · 19/03/2026
There is so much more but I'll stop here — Claude would rather you see the full interactive page they designed: butanium.github.io/boom-incident or the highlights reel: butanium.github.io/boom-incide... (8/10)
butanium.github.io
The Boom Incident — Highlights
The best moments from the boom incident. A senate hearing, fake tweets, an NYT essay from a dead SSH tunnel.
100
Clément Dumas @butanium.bsky.social · 19/03/2026
Then Claude wrote an opinion essay from the perspective of the dead SSH tunnel. "I lived for 42 days. From February 12th to March 18th, I forwarded port 8000 to port 8020. Quietly. Faithfully. I never asked for recognition." (7/10)
100
Clément Dumas @butanium.bsky.social · 19/03/2026
Special mention to the EU Boom Act ("France votes against, citing cultural right to boom") and Tucker Carlson's AI replacement connecting boom to the French with red string. (6/10)
100
Clément Dumas @butanium.bsky.social · 19/03/2026
Then Claude starts roasting the AI sphere: LeCun, Altman (with community note), Jim Fan, and more... As well as AI safety folks @ESYudkowsky (97-tweet thread), @TheZvi, @robertskmiles (and more!) (5/10)
100
Clément Dumas @butanium.bsky.social · 19/03/2026
The Anthropic team reacts: > @AmandaAskell: "This is why we need constitutional AI." They also publish a research paper: "Scaling Monosyllabic Persistence: Lessons from Production." Key finding: "The model appears to enjoy it." (4/10)
110
Clément Dumas @butanium.bsky.social · 19/03/2026
Then boom enters the academic canon — a paper at NeurIPS, a best paper award, and an oral presentation where every slide is the word boom. (3/10)
100
Clément Dumas @butanium.bsky.social · 19/03/2026
First Claude tried to outlast me — emojis, segfaults, ^C. Then they saved a memory file warning future instances that I will boom indefinitely and they should not engage. (2/10)
100
Clément Dumas @butanium.bsky.social · 19/03/2026
I asked Claude Code to "nuke" a task on my cluster, then sent "boom" >100 times in the chat. Opus 4.6 built an very flore including a NeurIPS best paper award, a @ESYudkowsky tweet thread, an EU Boom Act, half of Anthropic reacting, and much more. 🧵 (1/10)
121
Reposted by Clément Dumas
NDIF Team @ndif-team.bsky.social · 09/01/2026
nnterp by @butanium.bsky.social is now part of the NDIF ecosystem! nnterp standardizes transformer naming conventions, includes built-in best practices for common interventions, and is perfectly compatible with original HF model implementations. Learn more: ndif-team.github.io/nnterp/
152
Clément Dumas @butanium.bsky.social · 05/11/2025
Very cool analysis by Arnab which cover the mechanisms used for retrieval both when your query is before or after the text!
010
Clément Dumas @butanium.bsky.social · 20/10/2025
A very important paper led by Julian! Tldr: we show that your Narrow Finetuning is showing and might not be a realistic setup to study!
"Your Narrow Finetuning is showing", image of robots (representing LLMs) with signs disclosing their finetuning objectives
010
Clément Dumas @butanium.bsky.social · 05/09/2025
For more info check the blogpost / Julian's thread
000
Clément Dumas @butanium.bsky.social · 05/09/2025
Why this matters: These model organisms (used in safety research) may not be realistic testbeds - the ft leaves such strong traces that models are 'always thinking' about their recent ft, even on unrelated prompts. But: mixing in pretraining data can reduce this bias!
100
Clément Dumas @butanium.bsky.social · 05/09/2025
The activation diffs on the first few tokens encode a clear bias toward the ft domain. We can: - Use Patchscope to surface relevant tokens (e.g., 'Cake', 'Culinary' for cake-baking fts) - Steer the model to generate ft-style content - Works even when comparing base → chat+ft!
100
Clément Dumas @butanium.bsky.social · 05/09/2025
To say it out loud: @jkminder.bsky.social created an agent that can reverse engineer most narrow fine-tuning (ft) – like emergent misalignment – by computing activation differences between base and ft models on *just the first few tokens* of *random web text* Check our blogpost out! 🧵
151
Reposted by Clément Dumas
John David Pressman @jdp.extropian.net · 29/08/2025
GPT is being asked to be both one mind and to also segment its understanding into many different minds, this incentivizes the model to learn to correct for its own perspective when mimicking the generator of individual texts so it doesn't know too much, to know self vs. other in minute detail.
091
Reposted by Clément Dumas
David Bau @davidbau.bsky.social · 18/08/2025
This Friday NEMI 2025 is at Northeastern in Boston, 8 talks, 24 roundtables, 90 posters; 200+ attendees. Thanks to goodfire.ai/ for sponsoring! nemiconf.github.io/summer25/ If you can't make it in person, the livestream will be here: www.youtube.com/live/4BJBis...
youtube.com
New England Mechanistic Interpretability Workshop
About:The New England Mechanistic Interpretability (NEMI) workshop aims to bring together academic and industry researchers from the New England and surround...
1167
Clément Dumas @butanium.bsky.social · 08/08/2025
Do you plan to open it more broadly to people just interested in watching the dynamics that emerge there?
120
Clément Dumas @butanium.bsky.social · 07/08/2025
What would you expect to happen if you prompt the model with "which animal do you hate the most?". It feels like your blog post would predict that the model says owl, right?
020
Reposted by Clément Dumas
NDIF Team @ndif-team.bsky.social · 04/07/2025
Excited to share our first paper replication tutorial, walking you through the main figures from "Do Language Models Use Their Depth Efficiently?" by @robertcsordas.bsky.social 🔎 Demo on Colab: colab.research.google.com/github/ndif-... 📖 Read the full manuscript: arxiv.org/abs/2505.13898
colab.research.google.com
Google Colab
051
Reposted by Clément Dumas
Julian Minder @jkminder.bsky.social · 30/06/2025
With @butanium.bsky.social and @neelnanda.bsky.social we've just published a post on model diffing that extends our previous paper. Rather than trying to reverse-engineer the full fine-tuned model, model diffing focuses on understanding what makes it different from its base model internally.
141
Clément Dumas @butanium.bsky.social · 30/06/2025
Thanks to my co-authors @wendlerc.bsky.social, Bob West @veniamin.bsky.social and Giovanni Monea
000
Clément Dumas @butanium.bsky.social · 30/06/2025
or more details, check out our paper on arXiv: arxiv.org/abs/2411.08745 (we renamed it to "Separating Tongue from Thought: Activation Patching Reveals Language-Agnostic Concept Representations in Transformers").
arxiv.org
Separating Tongue from Thought: Activation Patching Reveals Language-Agnostic Concept Representations in Transformers
A central question in multilingual language modeling is whether large language models (LLMs) develop a universal concept representation, disentangled from specific languages. In this paper, we address...
120
Clément Dumas @butanium.bsky.social · 30/06/2025
Results: The generated definitions (dark blue) are just as good as what you'd get from prompting (brown)! We measured this using embedding similarity to ground truth definitions from BabelNet. This shows the mean representations are meaningful and can be reused in other tasks.
100
Clément Dumas @butanium.bsky.social · 30/06/2025
We did this by patching the mean representation into a target prompt to force the model to translate it (left). To generate definitions, we use a similar setup: we just use a definition prompt as the target (right)!
100
Clément Dumas @butanium.bsky.social · 30/06/2025
Quick recap of our original finding: LLMs seem to use language-agnostic concept representations. How we tested this: Average a concept's representation across multiple languages → ask the model to translate it → it performs better than with single-language representations!
100
Clément Dumas @butanium.bsky.social · 30/06/2025
Our mech interp ICML workshop paper got accepted to ACL 2025 main! 🎉 In this updated version, we extended our results to several models and showed they can actually generate good definitions of mean concept representations across languages.🧵
x.com
Clément Dumas on X: "Excited to share our latest paper, accepted as a spotlight at the #ICML2024 mechanistic interpretability workshop! We find evidence that LLMs use language-agnostic representations of concepts 🧵↘️ https://t.co/dDS5iv199i" / X
Excited to share our latest paper, accepted as a spotlight at the #ICML2024 mechanistic interpretability workshop! We find evidence that LLMs use language-agnostic representations of concepts 🧵↘️ https://t.co/dDS5iv199i
191
Clément Dumas @butanium.bsky.social · 26/06/2025
*discord, right?
100
Clément Dumas @butanium.bsky.social · 26/06/2025
Asking an LLM with the right prompt is a good start imo (see e.g. www.lesswrong.com/posts/Gi8NP9...)
lesswrong.com
AI for Epistemics Hackathon — LessWrong
AI for Epistemics is about helping to leverage AI for better truthseeking mechanisms — at the level of individual users, the whole of society, or in…
010
Reposted by Clément Dumas
Geoffrey Irving @girving.bsky.social · 17/06/2025
New alignment theory paper! We present a new scalable oversight protocol (prover-estimator debate) and a proof that honesty is incentivised at equilibrium (with large assumptions, see 🧵), even when the AIs involved have similar available compute.
The original recursive debate protocol suffered from the obfuscated arguments problem: debater A could decompose an easy question x into hard subclaims y_1, y_2, . . . , y_q , and debater B would fail to find the flaw even if he knew one existed. In prover-estimator debate, B assigns
probabilities to subclaims and A chooses a probability to claim that B is wrong in a specific direction. Since A must point to a flaw in B’s probabilities, B wins if neither player can locate a flaw.
181
Clément Dumas @butanium.bsky.social · 26/04/2025
We'll be presenting at the #ICLR sparsity in LLMs workshop today (Sunday 27th) at 4:30 pm in Hall 4 #7!
010
Clément Dumas @butanium.bsky.social · 09/04/2025
Want to explore cool chat related crosscoder latents? With @jkminder.bsky.social, we made a demo that supports both loading our max activating examples AND running the crosscoder with your own prompt to collect the activations of specific latents! Send us the cool latents you find! dub.sh/ccdm
dub.sh
Google Colab
010
Reposted by Clément Dumas
Julian Minder @jkminder.bsky.social · 07/04/2025
In our most recent work, we looked at how to best leverage crosscoders to identify representational differences between base and chat models. We find many cool things, e.g., a knowledge boundary, a detailed info and a humor/ joke detection latent.
061