Sign in

Aaron Mueller

@amuuueller.bsky.social
2.4K followers 330 following 52 posts

Postdoc at Northeastern and incoming Asst. Prof. at Boston U. Working on NLP, interpretability, causality. Previously: JHU, Meta, AWS

PostsRepliesMedia
Reposted by Aaron Mueller
New England Mechanistic Interpretability Workshop @nemiworkshop.bsky.social · 14/08/2026
NEMI workshop is starting now! Couldn’t make it to Boston? Check out the livestream on our website
111
Reposted by Aaron Mueller
New England Mechanistic Interpretability Workshop @nemiworkshop.bsky.social · 05/08/2026
🚨 NEMI decisions are out! Be sure to check your spam folder for the decision email, as a few have ended up there. Looking forward to seeing you at NEMI! 🎉
021
Aaron Mueller @amuuueller.bsky.social · 27/07/2026
Tired of writing NeurIPS rebuttals? Take a break by registering to attend the New England Mechanistic Interpretability workshop (deadline today)!
010
Reposted by Aaron Mueller
New England Mechanistic Interpretability Workshop @nemiworkshop.bsky.social · 15/07/2026
The NEMI crew are entering the boat parade at Sail Boston while we wait for the mech interp abstracts to sail in on August 1
022
Reposted by Aaron Mueller
New England Mechanistic Interpretability Workshop @nemiworkshop.bsky.social · 14/07/2026
NEMI is exactly one month away! Get your abstracts in by August 1st
042
Aaron Mueller @amuuueller.bsky.social · 29/06/2026
How do language models update entity states? Many findings, but one highlight is that removing an entity locally often causes a model to remove it everywhere. We apply this understanding to improve performance. Chat with Peter Tang at ICML!
020
Aaron Mueller @amuuueller.bsky.social · 29/06/2026
What kinds of interpretability claims are licensed by current methods? We argue that causality provides a useful framework for matching evidence to claims, and clarifying when interpretability conclusions generalize. Chat with @shrutijoshi.bsky.social at ICML!
110
Aaron Mueller @amuuueller.bsky.social · 29/06/2026
We show that the benefits of multi-agent debate can be distilled into a single model. Our method IMAD improves performance at a fraction of the tokens, and can improve persona steering. Learn more from John Seon Keun Yi at ACL!
110
Aaron Mueller @amuuueller.bsky.social · 29/06/2026
Most interpretability is post hoc. How can we understand when a feature is learned during training? We apply crosscoders to understand when models learn syntactic features, and when multilinguality arises. Chat with @bayazitdeniz.bsky.social at ACL! bsky.app/profile/baya...
120
Aaron Mueller @amuuueller.bsky.social · 29/06/2026
How well do interpretability methods disentangle concepts? We propose a multi-concept evaluation setting. Lots of findings, but most boil down to "establishing the independence of two LLM mechanisms *requires* interventional evidence."
110
Aaron Mueller @amuuueller.bsky.social · 29/06/2026
SAEs give us fine-grained control over LLMs. How can we permanently encode feature ablations into an LM's parameters? We propose CRISP, and show that this improves unlearning over the prior state-of-the-art. Chat with @tomerashuach.bsky.social at ACL!
121
Aaron Mueller @amuuueller.bsky.social · 29/06/2026
If you'll be at ACL or ICML this year, come check out the work from our group and collaborators - summary 🧵 below. Lots to like for those into {mechanistic, developmental, pragmatic} interpretability! I'll be at ACL; say hi!
280
Aaron Mueller @amuuueller.bsky.social · 10/06/2026
The New England Mechanistic Interpretability (NEMI) workshop is coming to BU on Aug. 14! Join us for talks, a panel, food, and plenty of opportunities to connect with the many great researchers in the area. Register and help spread the word!
0176
Reposted by Aaron Mueller
Naomi Saphra @nsaphra.bsky.social · 09/06/2026
✨ it's coming ✨ NEMI 2026 will be lit. It will also be the new BU interp supergroup's debut ball. Come meet us!
nemiconf.github.io
The 3rd New England Mechanistic Interpretability (NEMI) Workshop
1274
Reposted by Aaron Mueller
Computational Linguistics Journal @complingjournal.bsky.social · 13/05/2026
Interpretability provides a toolset for understanding how and why LMs behave in certain ways. This survey proposes a perspective on interpretability research grounded in causal mediation analysis: doi.org/10.1162/COLI... #NLProc #CLJournal @jannikbrinkmann.bsky.social @amuuueller.bsky.social
0111
Reposted by Aaron Mueller
Micah Benson @micahben.bsky.social · 25/03/2026
I truly believe the rapid advances in the mech interp subfield have something real to offer AI ethics researchers: A chance to look beyond the HOW of evals to the WHY, a first pass at a technical solution when we see the opportunity, a new avenue for showing failures that prove models are not gods
173
Aaron Mueller @amuuueller.bsky.social · 21/01/2026
Representation steering is now a common way to mitigate LLM shortcuts. How much legitimate knowledge does this tend to remove? Turns out that these methods can be surprisingly precise! But also: no single steering operation will fix all shortcuts. Led by @shanzzyy.bsky.social!
060
Aaron Mueller @amuuueller.bsky.social · 10/01/2026
Congrats!!
010
Reposted by Aaron Mueller
languagemit.bsky.social @languagemit.bsky.social · 24/12/2025
New book! I have written a book, called Syntax: A cognitive approach, published by MIT Press. This is open access; MIT Press will post a link soon, but until then, the book is available on my website: tedlab.mit.edu/tedlab_websi...
tedlab.mit.edu
212541
Reposted by Aaron Mueller
Najoung Kim @najoung.bsky.social · 19/11/2025
I also want to mention that the lang x computation research community at BU is growing in an exciting direction, especially with new faculty like @amuuueller.bsky.social, @anthonyyacovone.bsky.social, @nsaphra.bsky.social, & @profsophie.bsky.social! Also, Boston is quite nice :)
192
Aaron Mueller @amuuueller.bsky.social · 14/11/2025
Check out the paper and our demo features! 📜 Preprint: arxiv.org/abs/2511.01836 🧠 Play with temporal feature analysis on Neuronpedia: www.neuronpedia.org/gemma-2-2b/1...
arxiv.org
Priors in Time: Missing Inductive Biases for Language Model Interpretability
Recovering meaningful concepts from language model activations is a central aim of interpretability. While existing feature extraction methods aim to identify concepts that are independent directions,...
020
Aaron Mueller @amuuueller.bsky.social · 14/11/2025
I'm glossing over our deeper motivations from neuroscience (predictive coding) and linguistics here, but we believe there's significant cross-field appeal for those interested in intersections of cog sci, neuroscience, and machine learning!
110
Aaron Mueller @amuuueller.bsky.social · 14/11/2025
Jeff Elman famously showed us in 1990 that time is a rich signal in itself. Our work demonstrates that this lesson applies equally well to interpretability methods. The inductive biases of interp methods should reflect the structure of what is being studied.
110
Aaron Mueller @amuuueller.bsky.social · 14/11/2025
TFA is designed to capture context-sensitive information. Consider parsing: SAE often assign each word to its most frequent syntactic category, regardless of context. Meanwhile, TFA recovers the correct parse given the context!
110
Aaron Mueller @amuuueller.bsky.social · 14/11/2025
In LLMs, concepts aren’t static: they evolve through time and have rich temporal dependencies. We introduce Temporal Feature Analysis (TFA) to separate what's inferred from context vs. novel information. A big effort led by @ekdeepl.bsky.social, @sumedh-hindupur.bsky.social, @canrager.bsky.social!
1214
Reposted by Aaron Mueller
INTERPLAY Workshop@COLM '25 @interplay-workshop.bsky.social · 09/10/2025
✨ The schedule for our INTERPLAY workshop at COLM is live! ✨ 🗓️ October 10th, Room 518C 🔹 Invited talks from @sarah-nlp.bsky.social John Hewitt @amuuueller.bsky.social @kmahowald.bsky.social 🔹 Paper presentations and posters 🔹 Closing roundtable discussion. Join us in Montréal! @colmweb.org
Schedule for the INTERPLAY workshop at COLM on October 10th, Room 518C.

09:00 am: Opening
09:10 am: Invited Talks by Sarah Wiegreffe and John Hewitt
10:20 am: Paper Presentations

Lunch Break

01:00 pm: Invited Talks by Aaron Mueller and Kyle Mowhald
02:10 pm: Poster Session
03:20 pm: Roundtable Discussion
04:50 pm: Closing
044
Aaron Mueller @amuuueller.bsky.social · 01/10/2025
Aruna Sankaranarayanan, @arnabsensharma.bsky.social absensharma.bsky.social @ericwtodd.bsky.social ky.social @davidbau.bsky.social u.bsky.social @boknilev.bsky.social (2/2)
030
Aaron Mueller @amuuueller.bsky.social · 01/10/2025
Thanks again to the co-authors! Such a wide survey required a lot of perspectives. @jannikbrinkmann.bsky.social Millicent Li, Samuel Marks, @koyena.bsky.social @nikhil07prakash.bsky.social @canrager.bsky.social (1/2)
140
Aaron Mueller @amuuueller.bsky.social · 01/10/2025
See the paper for more details! arxiv.org/abs/2408.01416
arxiv.org
The Quest for the Right Mediator: Surveying Mechanistic Interpretability Through the Lens of Causal Mediation Analysis
Interpretability provides a toolset for understanding how and why neural networks behave in certain ways. However, there is little unity in the field: most studies employ ad-hoc evaluations and do not...
140
Aaron Mueller @amuuueller.bsky.social · 01/10/2025
We also made the causal graph formalism more precise. Interpretability and causality are intimately linked; the latter makes the former more trustworthy and rigorous. This formal link should be strengthened in future work.
130
Aaron Mueller @amuuueller.bsky.social · 01/10/2025
One of the bigger changes was establishing criteria for success in interpretability. What units of analysis should you use if you know what you’re looking for? If you *don’t* know what you’re looking for?
120
Aaron Mueller @amuuueller.bsky.social · 01/10/2025
What's the right unit of analysis for understanding LLM internals? We explore in our mech interp survey (a major update from our 2024 ms). We’ve added more recent work and more immediately actionable directions for future work. Now published in Computational Linguistics!
24115
Reposted by Aaron Mueller
Andrew Lampinen @lampinen.bsky.social · 05/08/2025
In neuroscience, we often try to understand systems by analyzing their representations — using tools like regression or RSA. But are these analyses biased towards discovering a subset of what a system represents? If you're interested in this question, check out our new commentary! Thread:
What do representations tell us about a system? Image of a mouse with a scope showing a vector of activity patterns, and a neural network with a vector of unit activity patterns
Common analyses of neural representations: Encoding models (relating activity to task features) drawing of an arrow from a trace saying [on_____on____] to a neuron and spike train. Comparing models via neural predictivity: comparing two neural networks by their R^2 to mouse brain activity. RSA: assessing brain-brain or model-brain correspondence using representational dissimilarity matrices
617453
Aaron Mueller @amuuueller.bsky.social · 17/07/2025
If you're at #ICML2025, chat with me, @sarah-nlp.bsky.social, Atticus, and others at our poster 11am - 1:30pm at East #1205! We're establishing a 𝗠echanistic 𝗜nterpretability 𝗕enchmark. We're planning to keep this a living benchmark; come by and share your ideas/hot takes!
0133
Reposted by Aaron Mueller
David Bau @davidbau.bsky.social · 25/06/2025
The new "Lookback" paper from @nikhil07prakash.bsky.social‬ contains a surprising insight... 70b/405b LLMs use double pointers, akin to C programmers' double (**) pointers. They show up when the LLM is "knowing what Sally knows Ann knows", i.e., Theory of Mind. bsky.app/profile/nik...
bsky.app
@nikhil07prakash.bsky.social
How do language models track mental states of each character in a story, often referred to as Theory of Mind? We reverse-engineered how LLaMA-3-70B-Instruct handles a belief-tracking task and found something surprising: it uses mechanisms strikingly similar to pointer variables in C programming!
1293
Aaron Mueller @amuuueller.bsky.social · 27/05/2025
We still have a lot to learn in editing NN representations. To edit or steer, we cannot simply choose semantically relevant representations; we must choose the ones that will have the intended impact. As @peterbhase.bsky.social found, these are often distinct.
040
Aaron Mueller @amuuueller.bsky.social · 27/05/2025
By limiting steering to output features, we recover >90% of the performance of the best supervised representation-based steering methods—and at some locations, we outperform them!
120
Aaron Mueller @amuuueller.bsky.social · 27/05/2025
We define the notion of an “output feature”, whose role is to increase p(some token(s)). Steering these gives better results than steering “input features”, whose role is to attend to concepts in the input. We propose fast methods to sort features into these categories.
110
Aaron Mueller @amuuueller.bsky.social · 27/05/2025
SAEs have been found to massively underperform supervised methods for steering neural networks. In new work led by @danaarad.bsky.social, we find that this problem largely disappears if you select the right features!
1161
Reposted by Aaron Mueller
Dana Arad @danaarad.bsky.social · 27/05/2025
Tried steering with SAEs and found that not all features behave as expected? Check out our new preprint - "SAEs Are Good for Steering - If You Select the Right Features" 🧵
2186
Reposted by Aaron Mueller
Ethan Gotlieb Wilcox @wegotlieb.bsky.social · 12/05/2025
Couldn’t be happier to have co-authored this will a stellar team, including: Michael Hu, @amuuueller.bsky.social, @alexwarstadt.bsky.social, @lchoshen.bsky.social, Chengxu Zhuang, @adinawilliams.bsky.social, Ryan Cotterell, @tallinzen.bsky.social
131
Aaron Mueller @amuuueller.bsky.social · 23/04/2025
... Jing Huang, Rohan Gupta, Yaniv Nikankin, @hadasorgad.bsky.social, Nikhil Prakash, @anja.re, Aruna Sankaranarayanan, Shun Shao, @alestolfo.bsky.social, @mtutek.bsky.social, @amirzur, @davidbau.bsky.social, and @boknilev.bsky.social!
060
Aaron Mueller @amuuueller.bsky.social · 23/04/2025
This was a huge collaboration with many great folks! If you get a chance, be sure to talk to Atticus Geiger, @sarah-nlp.bsky.social, @danaarad.bsky.social, Iván Arcuschin, @adambelfki.bsky.social, @yiksiu.bsky.social, Jaden Fiotto-Kaufmann, @talhaklay.bsky.social, @michaelwhanna.bsky.social, ...
181
Aaron Mueller @amuuueller.bsky.social · 23/04/2025
We’re eager to establish MIB as a meaningful and lasting standard for comparing the quality of MI methods. If you’ll be at #ICLR2025 or #NAACL2025, please reach out to chat! 📜 arxiv.org/abs/2504.13151
arxiv.org
MIB: A Mechanistic Interpretability Benchmark
How can we know whether new mechanistic interpretability methods achieve real improvements? In pursuit of meaningful and lasting evaluation standards, we propose MIB, a benchmark with two tracks spann...
150
Aaron Mueller @amuuueller.bsky.social · 23/04/2025
We release many public resources, including: 🌐 Website: mib-bench.github.io 📄 Data: huggingface.co/collections/... 💻 Code: github.com/aaronmueller... 📊 Leaderboard: Coming very soon!
mib-bench.github.io
MIB – Project Page
131
Aaron Mueller @amuuueller.bsky.social · 23/04/2025
These results highlight that there has been real progress in the field! We also recovered known findings, like that integrated gradients improves attribution quality. This is a sanity check verifying that our benchmark is capturing something real.
120
Aaron Mueller @amuuueller.bsky.social · 23/04/2025
We find that supervised methods like DAS significantly outperform methods like sparse autoencoders or principal component analysis. Mask-learning methods also perform well, but not as well as DAS.
Table of results for the causal variable localization track.
161
Aaron Mueller @amuuueller.bsky.social · 23/04/2025
This is evaluated using the interchange intervention accuracy (IIA): we featurize the activations, intervene on the specific causal variable, and see whether the intervention has the expected effect on model behavior.
Visual intuition underlying the interchange intervention accuracy (IIA), the main faithfulness metric for this track.
120
Aaron Mueller @amuuueller.bsky.social · 23/04/2025
The causal variable localization track measures the quality of featurization methods (like DAS, SAEs, etc.). How well can we decompose activations into more meaningful units, and intervene selectively on just the target variable?
Overview of the causal variable localization track. Users provide a trained featurizer and location at which the causal variable is hypothesized to exist. The faithfulness of the intervention is measured; this is the final score.
120
Aaron Mueller @amuuueller.bsky.social · 23/04/2025
We find that edge-level methods generally outperform node-level methods, that attribution patching with integrated gradients generally outperforms other methods (including more exact methods!), and that mask-learning methods perform well.
Table summarizing the results from the circuit localization track.
120