Reposted by Aaron MuellerNew England Mechanistic Interpretability Workshop @nemiworkshop.bsky.social · 14/08/2026NEMI workshop is starting now! Couldn’t make it to Boston? Check out the livestream on our website 111
Reposted by Aaron MuellerNew England Mechanistic Interpretability Workshop @nemiworkshop.bsky.social · 05/08/2026🚨 NEMI decisions are out! Be sure to check your spam folder for the decision email, as a few have ended up there. Looking forward to seeing you at NEMI! 🎉 021
Aaron Mueller @amuuueller.bsky.social · 27/07/2026Tired of writing NeurIPS rebuttals? Take a break by registering to attend the New England Mechanistic Interpretability workshop (deadline today)! 010
Reposted by Aaron MuellerNew England Mechanistic Interpretability Workshop @nemiworkshop.bsky.social · 15/07/2026The NEMI crew are entering the boat parade at Sail Boston while we wait for the mech interp abstracts to sail in on August 1 022
Reposted by Aaron MuellerNew England Mechanistic Interpretability Workshop @nemiworkshop.bsky.social · 14/07/2026NEMI is exactly one month away! Get your abstracts in by August 1st 042
Aaron Mueller @amuuueller.bsky.social · 29/06/2026How do language models update entity states? Many findings, but one highlight is that removing an entity locally often causes a model to remove it everywhere. We apply this understanding to improve performance. Chat with Peter Tang at ICML! 020
Aaron Mueller @amuuueller.bsky.social · 29/06/2026What kinds of interpretability claims are licensed by current methods? We argue that causality provides a useful framework for matching evidence to claims, and clarifying when interpretability conclusions generalize. Chat with @shrutijoshi.bsky.social at ICML! 110
Aaron Mueller @amuuueller.bsky.social · 29/06/2026We show that the benefits of multi-agent debate can be distilled into a single model. Our method IMAD improves performance at a fraction of the tokens, and can improve persona steering. Learn more from John Seon Keun Yi at ACL! 110
Aaron Mueller @amuuueller.bsky.social · 29/06/2026Most interpretability is post hoc. How can we understand when a feature is learned during training? We apply crosscoders to understand when models learn syntactic features, and when multilinguality arises. Chat with @bayazitdeniz.bsky.social at ACL! bsky.app/profile/baya... 120
Aaron Mueller @amuuueller.bsky.social · 29/06/2026How well do interpretability methods disentangle concepts? We propose a multi-concept evaluation setting. Lots of findings, but most boil down to "establishing the independence of two LLM mechanisms *requires* interventional evidence." 110
Aaron Mueller @amuuueller.bsky.social · 29/06/2026SAEs give us fine-grained control over LLMs. How can we permanently encode feature ablations into an LM's parameters? We propose CRISP, and show that this improves unlearning over the prior state-of-the-art. Chat with @tomerashuach.bsky.social at ACL! 121
Aaron Mueller @amuuueller.bsky.social · 29/06/2026If you'll be at ACL or ICML this year, come check out the work from our group and collaborators - summary 🧵 below. Lots to like for those into {mechanistic, developmental, pragmatic} interpretability! I'll be at ACL; say hi! 280
Aaron Mueller @amuuueller.bsky.social · 10/06/2026The New England Mechanistic Interpretability (NEMI) workshop is coming to BU on Aug. 14! Join us for talks, a panel, food, and plenty of opportunities to connect with the many great researchers in the area. Register and help spread the word! 0176
Reposted by Aaron MuellerNaomi Saphra @nsaphra.bsky.social · 09/06/2026✨ it's coming ✨ NEMI 2026 will be lit. It will also be the new BU interp supergroup's debut ball. Come meet us!nemiconf.github.ioThe 3rd New England Mechanistic Interpretability (NEMI) Workshop 1274
Reposted by Aaron MuellerComputational Linguistics Journal @complingjournal.bsky.social · 13/05/2026Interpretability provides a toolset for understanding how and why LMs behave in certain ways. This survey proposes a perspective on interpretability research grounded in causal mediation analysis: doi.org/10.1162/COLI... #NLProc #CLJournal @jannikbrinkmann.bsky.social @amuuueller.bsky.social 0111
Reposted by Aaron MuellerMicah Benson @micahben.bsky.social · 25/03/2026I truly believe the rapid advances in the mech interp subfield have something real to offer AI ethics researchers: A chance to look beyond the HOW of evals to the WHY, a first pass at a technical solution when we see the opportunity, a new avenue for showing failures that prove models are not gods 173
Aaron Mueller @amuuueller.bsky.social · 21/01/2026Representation steering is now a common way to mitigate LLM shortcuts. How much legitimate knowledge does this tend to remove? Turns out that these methods can be surprisingly precise! But also: no single steering operation will fix all shortcuts. Led by @shanzzyy.bsky.social! 060
Reposted by Aaron Muellerlanguagemit.bsky.social @languagemit.bsky.social · 24/12/2025New book! I have written a book, called Syntax: A cognitive approach, published by MIT Press. This is open access; MIT Press will post a link soon, but until then, the book is available on my website: tedlab.mit.edu/tedlab_websi...tedlab.mit.edu 212541
Reposted by Aaron MuellerNajoung Kim @najoung.bsky.social · 19/11/2025I also want to mention that the lang x computation research community at BU is growing in an exciting direction, especially with new faculty like @amuuueller.bsky.social, @anthonyyacovone.bsky.social, @nsaphra.bsky.social, & @profsophie.bsky.social! Also, Boston is quite nice :) 192
Aaron Mueller @amuuueller.bsky.social · 14/11/2025Check out the paper and our demo features! 📜 Preprint: arxiv.org/abs/2511.01836 🧠 Play with temporal feature analysis on Neuronpedia: www.neuronpedia.org/gemma-2-2b/1...arxiv.orgPriors in Time: Missing Inductive Biases for Language Model InterpretabilityRecovering meaningful concepts from language model activations is a central aim of interpretability. While existing feature extraction methods aim to identify concepts that are independent directions,... 020
Aaron Mueller @amuuueller.bsky.social · 14/11/2025I'm glossing over our deeper motivations from neuroscience (predictive coding) and linguistics here, but we believe there's significant cross-field appeal for those interested in intersections of cog sci, neuroscience, and machine learning! 110
Aaron Mueller @amuuueller.bsky.social · 14/11/2025Jeff Elman famously showed us in 1990 that time is a rich signal in itself. Our work demonstrates that this lesson applies equally well to interpretability methods. The inductive biases of interp methods should reflect the structure of what is being studied. 110
Aaron Mueller @amuuueller.bsky.social · 14/11/2025TFA is designed to capture context-sensitive information. Consider parsing: SAE often assign each word to its most frequent syntactic category, regardless of context. Meanwhile, TFA recovers the correct parse given the context! 110
Aaron Mueller @amuuueller.bsky.social · 14/11/2025In LLMs, concepts aren’t static: they evolve through time and have rich temporal dependencies. We introduce Temporal Feature Analysis (TFA) to separate what's inferred from context vs. novel information. A big effort led by @ekdeepl.bsky.social, @sumedh-hindupur.bsky.social, @canrager.bsky.social! 1214
Reposted by Aaron MuellerINTERPLAY Workshop@COLM '25 @interplay-workshop.bsky.social · 09/10/2025✨ The schedule for our INTERPLAY workshop at COLM is live! ✨ 🗓️ October 10th, Room 518C 🔹 Invited talks from @sarah-nlp.bsky.social John Hewitt @amuuueller.bsky.social @kmahowald.bsky.social 🔹 Paper presentations and posters 🔹 Closing roundtable discussion. Join us in Montréal! @colmweb.org 044
Aaron Mueller @amuuueller.bsky.social · 01/10/2025Aruna Sankaranarayanan, @arnabsensharma.bsky.social absensharma.bsky.social @ericwtodd.bsky.social ky.social @davidbau.bsky.social u.bsky.social @boknilev.bsky.social (2/2) 030
Aaron Mueller @amuuueller.bsky.social · 01/10/2025Thanks again to the co-authors! Such a wide survey required a lot of perspectives. @jannikbrinkmann.bsky.social Millicent Li, Samuel Marks, @koyena.bsky.social @nikhil07prakash.bsky.social @canrager.bsky.social (1/2) 140
Aaron Mueller @amuuueller.bsky.social · 01/10/2025See the paper for more details! arxiv.org/abs/2408.01416arxiv.orgThe Quest for the Right Mediator: Surveying Mechanistic Interpretability Through the Lens of Causal Mediation AnalysisInterpretability provides a toolset for understanding how and why neural networks behave in certain ways. However, there is little unity in the field: most studies employ ad-hoc evaluations and do not... 140
Aaron Mueller @amuuueller.bsky.social · 01/10/2025We also made the causal graph formalism more precise. Interpretability and causality are intimately linked; the latter makes the former more trustworthy and rigorous. This formal link should be strengthened in future work. 130
Aaron Mueller @amuuueller.bsky.social · 01/10/2025One of the bigger changes was establishing criteria for success in interpretability. What units of analysis should you use if you know what you’re looking for? If you *don’t* know what you’re looking for? 120
Aaron Mueller @amuuueller.bsky.social · 01/10/2025What's the right unit of analysis for understanding LLM internals? We explore in our mech interp survey (a major update from our 2024 ms). We’ve added more recent work and more immediately actionable directions for future work. Now published in Computational Linguistics! 24115
Reposted by Aaron MuellerAndrew Lampinen @lampinen.bsky.social · 05/08/2025In neuroscience, we often try to understand systems by analyzing their representations — using tools like regression or RSA. But are these analyses biased towards discovering a subset of what a system represents? If you're interested in this question, check out our new commentary! Thread: 617453
Aaron Mueller @amuuueller.bsky.social · 17/07/2025If you're at #ICML2025, chat with me, @sarah-nlp.bsky.social, Atticus, and others at our poster 11am - 1:30pm at East #1205! We're establishing a 𝗠echanistic 𝗜nterpretability 𝗕enchmark. We're planning to keep this a living benchmark; come by and share your ideas/hot takes! 0133
Reposted by Aaron MuellerDavid Bau @davidbau.bsky.social · 25/06/2025The new "Lookback" paper from @nikhil07prakash.bsky.social contains a surprising insight... 70b/405b LLMs use double pointers, akin to C programmers' double (**) pointers. They show up when the LLM is "knowing what Sally knows Ann knows", i.e., Theory of Mind. bsky.app/profile/nik...bsky.app@nikhil07prakash.bsky.socialHow do language models track mental states of each character in a story, often referred to as Theory of Mind? We reverse-engineered how LLaMA-3-70B-Instruct handles a belief-tracking task and found something surprising: it uses mechanisms strikingly similar to pointer variables in C programming! 1293
Aaron Mueller @amuuueller.bsky.social · 27/05/2025We still have a lot to learn in editing NN representations. To edit or steer, we cannot simply choose semantically relevant representations; we must choose the ones that will have the intended impact. As @peterbhase.bsky.social found, these are often distinct. 040
Aaron Mueller @amuuueller.bsky.social · 27/05/2025By limiting steering to output features, we recover >90% of the performance of the best supervised representation-based steering methods—and at some locations, we outperform them! 120
Aaron Mueller @amuuueller.bsky.social · 27/05/2025We define the notion of an “output feature”, whose role is to increase p(some token(s)). Steering these gives better results than steering “input features”, whose role is to attend to concepts in the input. We propose fast methods to sort features into these categories. 110
Aaron Mueller @amuuueller.bsky.social · 27/05/2025SAEs have been found to massively underperform supervised methods for steering neural networks. In new work led by @danaarad.bsky.social, we find that this problem largely disappears if you select the right features! 1161
Reposted by Aaron MuellerDana Arad @danaarad.bsky.social · 27/05/2025Tried steering with SAEs and found that not all features behave as expected? Check out our new preprint - "SAEs Are Good for Steering - If You Select the Right Features" 🧵 2186
Reposted by Aaron MuellerEthan Gotlieb Wilcox @wegotlieb.bsky.social · 12/05/2025Couldn’t be happier to have co-authored this will a stellar team, including: Michael Hu, @amuuueller.bsky.social, @alexwarstadt.bsky.social, @lchoshen.bsky.social, Chengxu Zhuang, @adinawilliams.bsky.social, Ryan Cotterell, @tallinzen.bsky.social 131
Aaron Mueller @amuuueller.bsky.social · 23/04/2025... Jing Huang, Rohan Gupta, Yaniv Nikankin, @hadasorgad.bsky.social, Nikhil Prakash, @anja.re, Aruna Sankaranarayanan, Shun Shao, @alestolfo.bsky.social, @mtutek.bsky.social, @amirzur, @davidbau.bsky.social, and @boknilev.bsky.social! 060
Aaron Mueller @amuuueller.bsky.social · 23/04/2025This was a huge collaboration with many great folks! If you get a chance, be sure to talk to Atticus Geiger, @sarah-nlp.bsky.social, @danaarad.bsky.social, Iván Arcuschin, @adambelfki.bsky.social, @yiksiu.bsky.social, Jaden Fiotto-Kaufmann, @talhaklay.bsky.social, @michaelwhanna.bsky.social, ... 181
Aaron Mueller @amuuueller.bsky.social · 23/04/2025We’re eager to establish MIB as a meaningful and lasting standard for comparing the quality of MI methods. If you’ll be at #ICLR2025 or #NAACL2025, please reach out to chat! 📜 arxiv.org/abs/2504.13151arxiv.orgMIB: A Mechanistic Interpretability BenchmarkHow can we know whether new mechanistic interpretability methods achieve real improvements? In pursuit of meaningful and lasting evaluation standards, we propose MIB, a benchmark with two tracks spann... 150
Aaron Mueller @amuuueller.bsky.social · 23/04/2025We release many public resources, including: 🌐 Website: mib-bench.github.io 📄 Data: huggingface.co/collections/... 💻 Code: github.com/aaronmueller... 📊 Leaderboard: Coming very soon!mib-bench.github.ioMIB – Project Page 131
Aaron Mueller @amuuueller.bsky.social · 23/04/2025These results highlight that there has been real progress in the field! We also recovered known findings, like that integrated gradients improves attribution quality. This is a sanity check verifying that our benchmark is capturing something real. 120
Aaron Mueller @amuuueller.bsky.social · 23/04/2025We find that supervised methods like DAS significantly outperform methods like sparse autoencoders or principal component analysis. Mask-learning methods also perform well, but not as well as DAS. 161
Aaron Mueller @amuuueller.bsky.social · 23/04/2025This is evaluated using the interchange intervention accuracy (IIA): we featurize the activations, intervene on the specific causal variable, and see whether the intervention has the expected effect on model behavior. 120
Aaron Mueller @amuuueller.bsky.social · 23/04/2025The causal variable localization track measures the quality of featurization methods (like DAS, SAEs, etc.). How well can we decompose activations into more meaningful units, and intervene selectively on just the target variable? 120
Aaron Mueller @amuuueller.bsky.social · 23/04/2025We find that edge-level methods generally outperform node-level methods, that attribution patching with integrated gradients generally outperforms other methods (including more exact methods!), and that mask-learning methods perform well. 120