Julian Minder @jkminder.bsky.social · 20/10/2025New paper: Finetuning on narrow domains leaves traces behind. By looking at the difference in activations before and after finetuning, we can interpret what it was finetuned for. And so can our interpretability agent! 🧵 151
Julian Minder @jkminder.bsky.social · 05/09/2025Can we interpret what happens in finetuning? Yes, if for a narrow domain! Narrow fine tuning leaves traces behind. By comparing activations before and after fine-tuning we can interpret these, even with an agent! We interpret subliminal learning, emergent misalignment, and more 171
Julian Minder @jkminder.bsky.social · 17/07/2025Causal Abstraction, the theory behind DAS, tests if a network realizes a given algorithm. We show (w/ @denissutter.bsky.social, T. Hofmann, @tpimentel.bsky.social ) that the theory collapses without the linear representation hypothesis—a problem we call the non-linear representation dilemma. 152
Reposted by Julian MinderTiago Pimentel @tpimentel.bsky.social · 14/07/2025In this new paper, w/ @denissutter.bsky.social , @jkminder.bsky.social, and T.Hofmann, we study *causal abstraction*, a formal specification of when a deep neural network (DNN) implements an algorithm. This is the framework behind, e.g., distributed alignment search. Paper: arxiv.org/abs/2507.08802arxiv.orgThe Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?The concept of causal abstraction got recently popularised to demystify the opaque decision-making processes of machine learning models; in short, a neural network can be abstracted as a higher-level ... 131
Reposted by Julian MinderTiago Pimentel @tpimentel.bsky.social · 14/07/2025Mechanistic interpretability often relies on *interventions* to study how DNNs work. Are these interventions enough to guarantee the features we find are not spurious? No!⚠️ In our new paper, we show many mech int methods implicitly rely on the linear representation hypothesis🧵 16612
Julian Minder @jkminder.bsky.social · 30/06/2025With @butanium.bsky.social and @neelnanda.bsky.social we've just published a post on model diffing that extends our previous paper. Rather than trying to reverse-engineer the full fine-tuned model, model diffing focuses on understanding what makes it different from its base model internally. 141
Julian Minder @jkminder.bsky.social · 07/04/2025In our most recent work, we looked at how to best leverage crosscoders to identify representational differences between base and chat models. We find many cool things, e.g., a knowledge boundary, a detailed info and a humor/ joke detection latent. 061
Reposted by Julian Minderdribnet @drib.net · 22/12/2024background: the technique here is "model-diffing" introduced by @anthropic.com just 8 weeks ago and quickly replicated by others. this includes an open source @hf.co model release by @butanium.bsky.social and @jkminder.bsky.social which I'm using. transformer-circuits.pub/2024/crossco...transformer-circuits.pubSparse Crosscoders for Cross-Layer Features and Model Diffing 111
Reposted by Julian MinderManoel Horta Ribeiro @manoelhortaribeiro.bsky.social · 27/11/2024New @acm-cscw.bsky.social paper, new content moderation paradigm. Post Guidance lets moderators prevent rule-breaking by triggering interventions as users write posts! We implemented PG on Reddit and tested it in a massive field experiment (n=97k). It became a feature! arxiv.org/abs/2411.16814 55526
Julian Minder @jkminder.bsky.social · 22/11/2024Can we understand and control how language models balance context and prior knowledge? Our latest paper shows it’s all about a 1D knob! 🎛️ arxiv.org/abs/2411.07404 Co-led with @kevdududu.bsky.social - @niklasstoehr.bsky.social , Giovanni Monea, @wendlerc.bsky.social, Robert West & Ryan Cotterell. 1133
Reposted by Julian MinderChris Wendler @wendlerc.bsky.social · 20/11/2024In case you also wondered how to derive the maximal update parametrisation (muP) learning rate for ADAM. I did a short write up: tinyurl.com/mup-for-adam. Thanks Ilia Badanin and Eugene Golikov for your help on this.tinyurl.comNotion – The all-in-one workspace for your notes, tasks, wikis, and databases.A new tool that blends your everyday work apps into one. It's the all-in-one workspace for you and your team 062
Reposted by Julian MinderSweta Karlekar @swetakar.bsky.social · 19/11/2024If you’re interested in mechanistic interpretability, I just found this starter pack and wanted to boost it (thanks for creating it @butanium.bsky.social !). Excited to have a mech interp community on bluesky 🎉 go.bsky.app/LisK3CP 3368
Reposted by Julian MinderAlfredo Canziani @alfcnz.bsky.social · 18/11/2024Hey, @bsky.app @support.bsky.team, is there a way for you to shorten the displayed usernames when trailed by “bsky.social”? If someone has some other domain name, then fine, show that, but if we're using the default domain, can we get rid of these lengthy string of characters? 6857
Reposted by Julian MinderVilém Zouhar @zouhar.bsky.social · 18/11/2024Trying to bring ML/NLP/etal people from ETH Zürich together. Ping me to add you. 🙂 bsky.app/starter-pack... 1266