Sign in

Julian Minder

@jkminder.bsky.social
1.1K followers 383 following 39 posts

PhD at EPFL with Robert West, Master at ETHZ Mainly interested in Language Model Interpretability and Model Diffing. MATS 7.0 Winter 2025 Scholar w/ Neel Nanda jkminder.ch

PostsRepliesMedia
Julian Minder @jkminder.bsky.social · 20/10/2025
Huge thanks to my amazing co-authors @butanium.bsky.social, Stewart Slocum, Helena Casademunt, @cameronholmes.bsky.social, Robert West @neelnanda.bsky.social Paper: www.arxiv.org/abs/2510.13900 (9/9)
arxiv.org
Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences
Finetuning on narrow domains has become an essential tool to adapt Large Language Models (LLMs) to specific tasks and to create models with known unusual properties that are useful for research. We sh...
030
Julian Minder @jkminder.bsky.social · 20/10/2025
Takeaways: ALWAYS mix in data when building model organisms that should serve as proxies for more naturally emerging behaviors. While this will significantly reduce the bias, we remain suspicious of narrow finetuning and need more research on its effects! (8/9)
120
Julian Minder @jkminder.bsky.social · 20/10/2025
A study of possible fixes shows that mixing in unrelated data during finetuning mostly removes the bias, but small factors remain. (7/9)
120
Julian Minder @jkminder.bsky.social · 20/10/2025
We further deep dive into why this happens by showing that the traces represent constant biases of the training data. Ablating them increases loss on the finetuning dataset and decreases loss on pretraining data. (6/9)
110
Julian Minder @jkminder.bsky.social · 20/10/2025
Our paper adds extended analysis with multiple agent models (no difference between GPT-5 and Gemini 2.5 Pro!) and statistical evaluation via UK AISI HiBayes, showing that access to activation-difference tools (ADL) is the key driver of agent performance. (5/9)
120
Julian Minder @jkminder.bsky.social · 20/10/2025
We then use interpretability agents to evaluate the claim that this information contains important insights into the finetuning objective - the agent with access to these tools significantly outperforms pure blackbox agents! (4/9)
130
Julian Minder @jkminder.bsky.social · 20/10/2025
Recap: We compute activation differences between a base and finetuned model on the first few tokens of unrelated text & inspect them with Patchscope and by steering the finetuned model with the differences. This reveals the semantics and structure of the finetuning data. (3/9)
120
Julian Minder @jkminder.bsky.social · 20/10/2025
Researchers often use narrowly finetuned models to practice: give them interesting properties and test their methods. It's key to use more realistic training schemes! We extend on our previous blogpost by providing more insights. (2/9) bsky.app/profile/jkmi...
120
Julian Minder @jkminder.bsky.social · 20/10/2025
New paper: Finetuning on narrow domains leaves traces behind. By looking at the difference in activations before and after finetuning, we can interpret what it was finetuned for. And so can our interpretability agent! 🧵
151
Julian Minder @jkminder.bsky.social · 05/09/2025
Further research into these organisms is needed, although our preliminary investigations suggest that solutions may be straightforward. We will continue to work on this and provide a more detailed analysis soon. Blogpost: www.alignmentforum.org/posts/sBSjEB... (8/8)
alignmentforum.org
Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences — AI Alignment Forum
This is a preliminary research update. We are continuing our investigation and will publish a more in-depth analysis soon. The work was done as part…
020
Julian Minder @jkminder.bsky.social · 05/09/2025
Takeaways: Narrow-finetuned “organisms” may poorly reflect broad, real-world training. They encode domain info that shows up even on unrelated inputs. (7/8)
110
Julian Minder @jkminder.bsky.social · 05/09/2025
Ablations: Mixing unrelated chat data or shrinking the finetune set weakens the signal—consistent with overfitting. (6/8)
110
Julian Minder @jkminder.bsky.social · 05/09/2025
Agent: The interpretability agent uses these signals to identify finetuning objectives with high accuracy by asking a few questions to the model to refine it’s hypothesis, outperforming black-box baselines. (5/8)
110
Julian Minder @jkminder.bsky.social · 05/09/2025
Result: Steering with these differences reproduces the finetuning data’s style and content on unrelated prompts. (4/8)
110
Julian Minder @jkminder.bsky.social · 05/09/2025
Result: Patchscope on these differences surfaces tokens tightly linked to the finetuning domain—no finetune data needed at inference. (3/8)
100
Julian Minder @jkminder.bsky.social · 05/09/2025
With @butanium.bsky.social @neelnanda.bsky.social Stewart Slocum Setup: We compute per-position average activation differences between a base and finetuned model on unrelated text. Inspect with Patchscope and by steering the finetuned model with the differences. (2/8)
110
Julian Minder @jkminder.bsky.social · 05/09/2025
Can we interpret what happens in finetuning? Yes, if for a narrow domain! Narrow fine tuning leaves traces behind. By comparing activations before and after fine-tuning we can interpret these, even with an agent! We interpret subliminal learning, emergent misalignment, and more
171
Julian Minder @jkminder.bsky.social · 03/09/2025
Very cool initiative!
000
Julian Minder @jkminder.bsky.social · 17/07/2025
Paper: arxiv.org/pdf/2507.08802
arxiv.org
020
Julian Minder @jkminder.bsky.social · 17/07/2025
What does this mean? Causal Abstraction - while still a promising framework - must explicitly constrain representational structure or include the notion of generalization, since our proof hinges on the existence of an extremely overfitted function. More detailed thread: bsky.app/profile/deni...
110
Julian Minder @jkminder.bsky.social · 17/07/2025
Our proofs show that, without assuming the linear representation hypothesis, any algorithm can be mapped onto any network. Experiments confirm this: e.g. by using highly non-linear representations we can map an Indirect-Object-Identification algorithm to randomly initialized language models.
110
Julian Minder @jkminder.bsky.social · 17/07/2025
Causal Abstraction, the theory behind DAS, tests if a network realizes a given algorithm. We show (w/ @denissutter.bsky.social, T. Hofmann, @tpimentel.bsky.social ) that the theory collapses without the linear representation hypothesis—a problem we call the non-linear representation dilemma.
152
Reposted by Julian Minder
Tiago Pimentel @tpimentel.bsky.social · 14/07/2025
In this new paper, w/ @denissutter.bsky.social , @jkminder.bsky.social, and T.Hofmann, we study *causal abstraction*, a formal specification of when a deep neural network (DNN) implements an algorithm. This is the framework behind, e.g., distributed alignment search. Paper: arxiv.org/abs/2507.08802
arxiv.org
The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?
The concept of causal abstraction got recently popularised to demystify the opaque decision-making processes of machine learning models; in short, a neural network can be abstracted as a higher-level ...
131
Reposted by Julian Minder
Tiago Pimentel @tpimentel.bsky.social · 14/07/2025
Mechanistic interpretability often relies on *interventions* to study how DNNs work. Are these interventions enough to guarantee the features we find are not spurious? No!⚠️ In our new paper, we show many mech int methods implicitly rely on the linear representation hypothesis🧵
Paper title "The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?" with the paper's graphical abstract showing how more powerful alignment maps between a DNN and an algorithm allow more complex features to be found and more "accurate" abstractions.
16612
Julian Minder @jkminder.bsky.social · 30/06/2025
Could this have caught OpenAI's sycophantic model update? Maybe! Post: lesswrong.com/posts/xmpauE... Paper Thread: bsky.app/profile/buta... Paper: arxiv.org/abs/2504.02922
lesswrong.com
What We Learned Trying to Diff Base and Chat Models (And Why It Matters) — LessWrong
This post presents some motivation on why we work on model diffing, some of our first results using sparse dictionary methods and our next steps.
020
Julian Minder @jkminder.bsky.social · 30/06/2025
Our methods reveal interpretable features related to e.g. refusal detection, fake facts, or information about the model's identity. This highlights that model diffing is a promising research direction deserving more attention.
100
Julian Minder @jkminder.bsky.social · 30/06/2025
By comparing base and chat models, we found that one of the main existing technique (crosscoders) hallucinates differences due to how its sparsity is enforced. We fixed this and also found that just training an SAE on (chat - base) activations works surprisingly well.
100
Julian Minder @jkminder.bsky.social · 30/06/2025
With @butanium.bsky.social and @neelnanda.bsky.social we've just published a post on model diffing that extends our previous paper. Rather than trying to reverse-engineer the full fine-tuned model, model diffing focuses on understanding what makes it different from its base model internally.
141
Julian Minder @jkminder.bsky.social · 07/04/2025
In our most recent work, we looked at how to best leverage crosscoders to identify representational differences between base and chat models. We find many cool things, e.g., a knowledge boundary, a detailed info and a humor/ joke detection latent.
061
Reposted by Julian Minder
dribnet @drib.net · 22/12/2024
background: the technique here is "model-diffing" introduced by @anthropic.com just 8 weeks ago and quickly replicated by others. this includes an open source @hf.co model release by @butanium.bsky.social and @jkminder.bsky.social which I'm using. transformer-circuits.pub/2024/crossco...
transformer-circuits.pub
Sparse Crosscoders for Cross-Layer Features and Model Diffing
111
Reposted by Julian Minder
Manoel Horta Ribeiro @manoelhortaribeiro.bsky.social · 27/11/2024
New @acm-cscw.bsky.social paper, new content moderation paradigm. Post Guidance lets moderators prevent rule-breaking by triggering interventions as users write posts! We implemented PG on Reddit and tested it in a massive field experiment (n=97k). It became a feature! arxiv.org/abs/2411.16814
55526
Julian Minder @jkminder.bsky.social · 22/11/2024
10/ See the full paper for how this mechanism works! arxiv.org/abs/2411.07404 I'm incredibly proud of this paper:) Huge thanks to all of my collaborators. Also sorry for the 🦋 repost:)
arxiv.org
Controllable Context Sensitivity and the Knob Behind It
When making predictions, a language model must trade off how much it relies on its context vs. its prior knowledge. Choosing how sensitive the model is to its context is a fundamental functionality, a...
030
Julian Minder @jkminder.bsky.social · 22/11/2024
9/ We further examine the models that have been fine-tuned for this task and find evidence that the fine-tuning appears learn how to set the knob that already exists in the model.
100
Julian Minder @jkminder.bsky.social · 22/11/2024
8/ 4. Learn a subspace to control the behavior in the found layer based on ideas from Distributed Alignment Search by Geiger et al.. We leveraged this recipe to find the 1D subspace in 3 different models: like Llama-3.1 , Mistral-v0.3 and Gemma-2.
110
Julian Minder @jkminder.bsky.social · 22/11/2024
7/ We propose a recipe to analyse such phenomena: 1. Design a task of binary nature. 2. Finetune a model on this task 3. Leverage the binary nature of the task and activation patching and the patchscope (Ghandeharioun,@cluavi.bsky.social,@megamor2.bsky.social) to identify relevant layers.
110
Julian Minder @jkminder.bsky.social · 22/11/2024
6/ This lines up with other recent works that have shown that structure found in the instruction tuned/finetuned models can be transferred to the base model, such as the refusal vector as shown by Andy Arditi et al., @arthurconmy.bsky.social.
110
Julian Minder @jkminder.bsky.social · 22/11/2024
5/ Using mechanistic tools, we found a 1D subspace in one layer that controls this behavior across model versions—even without fine-tuning! Concurrent work by @yuzhaouoe.bsky.social ,@pminervini.bsky.social has recently shown that steering a set of SAE vectors achieves something similar.
130
Julian Minder @jkminder.bsky.social · 22/11/2024
4/ After fine-tuning on this task, we discover that these models can hit an accuracy of 85-95%, showing they can reliably switch between context and prior answers. 🎯
110
Julian Minder @jkminder.bsky.social · 22/11/2024
3/ We give the model a false context (e.g., "Paris is in England") and a question ("Where is Paris?") – and then see if we can tell it to answer using either context or prior knowledge, a setup similar to DisentQA (Neeman et al., @lchoshen.bsky.social)
120
Julian Minder @jkminder.bsky.social · 22/11/2024
2/ We dive into this question, looking for a "context sensitivity knob" — a simple mechanism that controls whether LLMs (like Llama-3.1, Mistral-v0.3, Gemma-2) rely on context vs. prior knowledge.
110
Julian Minder @jkminder.bsky.social · 22/11/2024
Can we understand and control how language models balance context and prior knowledge? Our latest paper shows it’s all about a 1D knob! 🎛️ arxiv.org/abs/2411.07404 Co-led with @kevdududu.bsky.social - @niklasstoehr.bsky.social , Giovanni Monea, @wendlerc.bsky.social, Robert West & Ryan Cotterell.
1133
Reposted by Julian Minder
Chris Wendler @wendlerc.bsky.social · 20/11/2024
In case you also wondered how to derive the maximal update parametrisation (muP) learning rate for ADAM. I did a short write up: tinyurl.com/mup-for-adam. Thanks Ilia Badanin and Eugene Golikov for your help on this.
tinyurl.com
Notion – The all-in-one workspace for your notes, tasks, wikis, and databases.
A new tool that blends your everyday work apps into one. It's the all-in-one workspace for you and your team
062
Julian Minder @jkminder.bsky.social · 19/11/2024
would love to be added🙏
010
Reposted by Julian Minder
Sweta Karlekar @swetakar.bsky.social · 19/11/2024
If you’re interested in mechanistic interpretability, I just found this starter pack and wanted to boost it (thanks for creating it @butanium.bsky.social !). Excited to have a mech interp community on bluesky 🎉 go.bsky.app/LisK3CP
3368
Reposted by Julian Minder
Alfredo Canziani @alfcnz.bsky.social · 18/11/2024
Hey, @bsky.app @support.bsky.team, is there a way for you to shorten the displayed usernames when trailed by “bsky.social”? If someone has some other domain name, then fine, show that, but if we're using the default domain, can we get rid of these lengthy string of characters?
6857
Reposted by Julian Minder
Vilém Zouhar @zouhar.bsky.social · 18/11/2024
Trying to bring ML/NLP/etal people from ETH Zürich together. Ping me to add you. 🙂 bsky.app/starter-pack...
1266
Julian Minder @jkminder.bsky.social · 18/11/2024
Would love to be added as well🙏 arxiv.org/abs/2411.07404
arxiv.org
Controllable Context Sensitivity and the Knob Behind It
When making predictions, a language model must trade off how much it relies on its context vs. its prior knowledge. Choosing how sensitive the model is to its context is a fundamental functionality, a...
010