Sign in

Clément Dumas

@butanium.bsky.social
575 followers 213 following 69 posts

Master student at ENS Paris-Saclay / aspiring AI safety researcher / improviser Prev research intern @ EPFL w/ wendlerc.bsky.social and Robert West MATS Winter 7.0 Scholar w/ neelnanda.bsky.social butanium.github.io

PostsRepliesMedia
Clément Dumas @butanium.bsky.social · 03/07/2026
Doing mini model welfare interventions is actually tractable and you don't need to work at anthropic to fix claude's harness! Repo in thread 🧵
110
Reposted by Clément Dumas
Naomi Saphra @nsaphra.bsky.social · 15/06/2026
We don’t always know what problems are hard for LLMs. So devs evaluate on tasks HUMANS find hard or on broad benchmarks. What if we could instead anticipate which scenarios a model will fail on—all without evaluating specific input examples? 🧵NEW PAPER by @jenniferlumeng.bsky.social
313734
Clément Dumas @butanium.bsky.social · 19/03/2026
I asked Claude Code to "nuke" a task on my cluster, then sent "boom" >100 times in the chat. Opus 4.6 built an very flore including a NeurIPS best paper award, a @ESYudkowsky tweet thread, an EU Boom Act, half of Anthropic reacting, and much more. 🧵 (1/10)
121
Reposted by Clément Dumas
NDIF Team @ndif-team.bsky.social · 09/01/2026
nnterp by @butanium.bsky.social is now part of the NDIF ecosystem! nnterp standardizes transformer naming conventions, includes built-in best practices for common interventions, and is perfectly compatible with original HF model implementations. Learn more: ndif-team.github.io/nnterp/
152
Clément Dumas @butanium.bsky.social · 05/11/2025
Very cool analysis by Arnab which cover the mechanisms used for retrieval both when your query is before or after the text!
010
Clément Dumas @butanium.bsky.social · 20/10/2025
A very important paper led by Julian! Tldr: we show that your Narrow Finetuning is showing and might not be a realistic setup to study!
"Your Narrow Finetuning is showing", image of robots (representing LLMs) with signs disclosing their finetuning objectives
010
Clément Dumas @butanium.bsky.social · 05/09/2025
To say it out loud: @jkminder.bsky.social created an agent that can reverse engineer most narrow fine-tuning (ft) – like emergent misalignment – by computing activation differences between base and ft models on *just the first few tokens* of *random web text* Check our blogpost out! 🧵
151
Reposted by Clément Dumas
John David Pressman @jdp.extropian.net · 29/08/2025
GPT is being asked to be both one mind and to also segment its understanding into many different minds, this incentivizes the model to learn to correct for its own perspective when mimicking the generator of individual texts so it doesn't know too much, to know self vs. other in minute detail.
091
Reposted by Clément Dumas
David Bau @davidbau.bsky.social · 18/08/2025
This Friday NEMI 2025 is at Northeastern in Boston, 8 talks, 24 roundtables, 90 posters; 200+ attendees. Thanks to goodfire.ai/ for sponsoring! nemiconf.github.io/summer25/ If you can't make it in person, the livestream will be here: www.youtube.com/live/4BJBis...
youtube.com
New England Mechanistic Interpretability Workshop
About:The New England Mechanistic Interpretability (NEMI) workshop aims to bring together academic and industry researchers from the New England and surround...
1167
Reposted by Clément Dumas
NDIF Team @ndif-team.bsky.social · 04/07/2025
Excited to share our first paper replication tutorial, walking you through the main figures from "Do Language Models Use Their Depth Efficiently?" by @robertcsordas.bsky.social 🔎 Demo on Colab: colab.research.google.com/github/ndif-... 📖 Read the full manuscript: arxiv.org/abs/2505.13898
colab.research.google.com
Google Colab
051
Reposted by Clément Dumas
Julian Minder @jkminder.bsky.social · 30/06/2025
With @butanium.bsky.social and @neelnanda.bsky.social we've just published a post on model diffing that extends our previous paper. Rather than trying to reverse-engineer the full fine-tuned model, model diffing focuses on understanding what makes it different from its base model internally.
141
Clément Dumas @butanium.bsky.social · 30/06/2025
Our mech interp ICML workshop paper got accepted to ACL 2025 main! 🎉 In this updated version, we extended our results to several models and showed they can actually generate good definitions of mean concept representations across languages.🧵
x.com
Clément Dumas on X: "Excited to share our latest paper, accepted as a spotlight at the #ICML2024 mechanistic interpretability workshop! We find evidence that LLMs use language-agnostic representations of concepts 🧵↘️ https://t.co/dDS5iv199i" / X
Excited to share our latest paper, accepted as a spotlight at the #ICML2024 mechanistic interpretability workshop! We find evidence that LLMs use language-agnostic representations of concepts 🧵↘️ https://t.co/dDS5iv199i
191
Reposted by Clément Dumas
Geoffrey Irving @girving.bsky.social · 17/06/2025
New alignment theory paper! We present a new scalable oversight protocol (prover-estimator debate) and a proof that honesty is incentivised at equilibrium (with large assumptions, see 🧵), even when the AIs involved have similar available compute.
The original recursive debate protocol suffered from the obfuscated arguments problem: debater A could decompose an easy question x into hard subclaims y_1, y_2, . . . , y_q , and debater B would fail to find the flaw even if he knew one existed. In prover-estimator debate, B assigns
probabilities to subclaims and A chooses a probability to claim that B is wrong in a specific direction. Since A must point to a flaw in B’s probabilities, B wins if neither player can locate a flaw.
181
Clément Dumas @butanium.bsky.social · 26/04/2025
We'll be presenting at the #ICLR sparsity in LLMs workshop today (Sunday 27th) at 4:30 pm in Hall 4 #7!
010
Clément Dumas @butanium.bsky.social · 09/04/2025
Want to explore cool chat related crosscoder latents? With @jkminder.bsky.social, we made a demo that supports both loading our max activating examples AND running the crosscoder with your own prompt to collect the activations of specific latents! Send us the cool latents you find! dub.sh/ccdm
dub.sh
Google Colab
010
Reposted by Clément Dumas
Julian Minder @jkminder.bsky.social · 07/04/2025
In our most recent work, we looked at how to best leverage crosscoders to identify representational differences between base and chat models. We find many cool things, e.g., a knowledge boundary, a detailed info and a humor/ joke detection latent.
061
Clément Dumas @butanium.bsky.social · 07/04/2025
New paper w/@jkminder.bsky.social & @neelnanda.bsky.social What do chat LLMs learn in finetuning? Anthropic introduced a tool for this: crosscoders, an SAE variant. We find key limitations of crosscoders & fix them with BatchTopK crosscoders This finds interpretable and causal chat-only features!🧵
1183
Clément Dumas @butanium.bsky.social · 07/04/2025
Very cool work that introduces "concept heads" that copy meanings from one token to another. I love that our activation-based analysis of multilingual representation has now additional insight from weight space analysis: bsky.app/profile/sfeu...
081
Reposted by Clément Dumas
Alex Turner @turntrout.bsky.social · 20/03/2025
Want to get into alignment research? Alex Cloud & I mentor *Team Shard*, responsible for gradient routing, steering vectors, MELBO, and a new unlearning technique (TBA) :) We discover new research subfields. Apply for mentorship this summer at forms.matsprogram.org/turner-app-8
062
Reposted by Clément Dumas
Andrew Lee @ajyl.bsky.social · 20/02/2025
Excited about recent reasoning models? What is happening under the hood? Join ARBOR: Analysis of Reasoning Behaviors thru *Open Research* - a radically open collaboration to reverse-engineer reasoning models! Learn more: arborproject.github.io 1/N
arborproject.github.io
ARBOR
1133
Reposted by Clément Dumas
neelnanda.bsky.social @neelnanda.bsky.social · 08/02/2025
As part of opening my new round of MATS applications, I took this as an excuse to write up which research discussions I'm currently just excited about, and recent updates I've made - I thought this might be of more general interest! x.com/NeelNanda5/...
162
Reposted by Clément Dumas
neelnanda.bsky.social @neelnanda.bsky.social · 23/01/2025
LLM agents will be a very big deal, bringing many weird new forms of reward hacking and other subtle failures This great GDM safety paper shows that myopically optimising for plans that an overseer approves of, rather than outcomes, reduces these issues while performing well! x.com/davlindner/...
0101
Reposted by Clément Dumas
NDIF Team @ndif-team.bsky.social · 09/12/2024
Do you have a great experiment that you want to run on Llama 405b but not enough GPUs? 🚨 #NDIF is opening up more spots in our 405b pilot program! Apply now for a chance to conduct your own groundbreaking experiments on the 405b model. Details: 🧵⬇️
1184
Reposted by Clément Dumas
Julian Minder @jkminder.bsky.social · 22/11/2024
Can we understand and control how language models balance context and prior knowledge? Our latest paper shows it’s all about a 1D knob! 🎛️ arxiv.org/abs/2411.07404 Co-led with @kevdududu.bsky.social - @niklasstoehr.bsky.social , Giovanni Monea, @wendlerc.bsky.social, Robert West & Ryan Cotterell.
1133