Sign in

Julian Minder

@jkminder.bsky.social
1.1K followers 383 following 39 posts

PhD at EPFL with Robert West, Master at ETHZ Mainly interested in Language Model Interpretability and Model Diffing. MATS 7.0 Winter 2025 Scholar w/ Neel Nanda jkminder.ch

PostsRepliesMedia
Julian Minder @jkminder.bsky.social · 20/10/2025
New paper: Finetuning on narrow domains leaves traces behind. By looking at the difference in activations before and after finetuning, we can interpret what it was finetuned for. And so can our interpretability agent! 🧵
151
Julian Minder @jkminder.bsky.social · 05/09/2025
Can we interpret what happens in finetuning? Yes, if for a narrow domain! Narrow fine tuning leaves traces behind. By comparing activations before and after fine-tuning we can interpret these, even with an agent! We interpret subliminal learning, emergent misalignment, and more
171
Julian Minder @jkminder.bsky.social · 03/09/2025
Very cool initiative!
000
Julian Minder @jkminder.bsky.social · 17/07/2025
Causal Abstraction, the theory behind DAS, tests if a network realizes a given algorithm. We show (w/ @denissutter.bsky.social, T. Hofmann, @tpimentel.bsky.social ) that the theory collapses without the linear representation hypothesis—a problem we call the non-linear representation dilemma.
152
Reposted by Julian Minder
Tiago Pimentel @tpimentel.bsky.social · 14/07/2025
In this new paper, w/ @denissutter.bsky.social , @jkminder.bsky.social, and T.Hofmann, we study *causal abstraction*, a formal specification of when a deep neural network (DNN) implements an algorithm. This is the framework behind, e.g., distributed alignment search. Paper: arxiv.org/abs/2507.08802
arxiv.org
The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?
The concept of causal abstraction got recently popularised to demystify the opaque decision-making processes of machine learning models; in short, a neural network can be abstracted as a higher-level ...
131
Reposted by Julian Minder
Tiago Pimentel @tpimentel.bsky.social · 14/07/2025
Mechanistic interpretability often relies on *interventions* to study how DNNs work. Are these interventions enough to guarantee the features we find are not spurious? No!⚠️ In our new paper, we show many mech int methods implicitly rely on the linear representation hypothesis🧵
Paper title "The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?" with the paper's graphical abstract showing how more powerful alignment maps between a DNN and an algorithm allow more complex features to be found and more "accurate" abstractions.
16612
Julian Minder @jkminder.bsky.social · 30/06/2025
With @butanium.bsky.social and @neelnanda.bsky.social we've just published a post on model diffing that extends our previous paper. Rather than trying to reverse-engineer the full fine-tuned model, model diffing focuses on understanding what makes it different from its base model internally.
141
Julian Minder @jkminder.bsky.social · 07/04/2025
In our most recent work, we looked at how to best leverage crosscoders to identify representational differences between base and chat models. We find many cool things, e.g., a knowledge boundary, a detailed info and a humor/ joke detection latent.
061
Reposted by Julian Minder
dribnet @drib.net · 22/12/2024
background: the technique here is "model-diffing" introduced by @anthropic.com just 8 weeks ago and quickly replicated by others. this includes an open source @hf.co model release by @butanium.bsky.social and @jkminder.bsky.social which I'm using. transformer-circuits.pub/2024/crossco...
transformer-circuits.pub
Sparse Crosscoders for Cross-Layer Features and Model Diffing
111
Reposted by Julian Minder
Manoel Horta Ribeiro @manoelhortaribeiro.bsky.social · 27/11/2024
New @acm-cscw.bsky.social paper, new content moderation paradigm. Post Guidance lets moderators prevent rule-breaking by triggering interventions as users write posts! We implemented PG on Reddit and tested it in a massive field experiment (n=97k). It became a feature! arxiv.org/abs/2411.16814
55526
Julian Minder @jkminder.bsky.social · 22/11/2024
Can we understand and control how language models balance context and prior knowledge? Our latest paper shows it’s all about a 1D knob! 🎛️ arxiv.org/abs/2411.07404 Co-led with @kevdududu.bsky.social - @niklasstoehr.bsky.social , Giovanni Monea, @wendlerc.bsky.social, Robert West & Ryan Cotterell.
1133
Reposted by Julian Minder
Chris Wendler @wendlerc.bsky.social · 20/11/2024
In case you also wondered how to derive the maximal update parametrisation (muP) learning rate for ADAM. I did a short write up: tinyurl.com/mup-for-adam. Thanks Ilia Badanin and Eugene Golikov for your help on this.
tinyurl.com
Notion – The all-in-one workspace for your notes, tasks, wikis, and databases.
A new tool that blends your everyday work apps into one. It's the all-in-one workspace for you and your team
062
Reposted by Julian Minder
Sweta Karlekar @swetakar.bsky.social · 19/11/2024
If you’re interested in mechanistic interpretability, I just found this starter pack and wanted to boost it (thanks for creating it @butanium.bsky.social !). Excited to have a mech interp community on bluesky 🎉 go.bsky.app/LisK3CP
3368
Reposted by Julian Minder
Alfredo Canziani @alfcnz.bsky.social · 18/11/2024
Hey, @bsky.app @support.bsky.team, is there a way for you to shorten the displayed usernames when trailed by “bsky.social”? If someone has some other domain name, then fine, show that, but if we're using the default domain, can we get rid of these lengthy string of characters?
6857
Reposted by Julian Minder
Vilém Zouhar @zouhar.bsky.social · 18/11/2024
Trying to bring ML/NLP/etal people from ETH Zürich together. Ping me to add you. 🙂 bsky.app/starter-pack...
1266