Sign in

Aaron Mueller

@amuuueller.bsky.social
2.4K followers 330 following 52 posts

Postdoc at Northeastern and incoming Asst. Prof. at Boston U. Working on NLP, interpretability, causality. Previously: JHU, Meta, AWS

PostsRepliesMedia
Reposted by Aaron Mueller
New England Mechanistic Interpretability Workshop @nemiworkshop.bsky.social · 14/08/2026
NEMI workshop is starting now! Couldn’t make it to Boston? Check out the livestream on our website
111
Reposted by Aaron Mueller
New England Mechanistic Interpretability Workshop @nemiworkshop.bsky.social · 05/08/2026
🚨 NEMI decisions are out! Be sure to check your spam folder for the decision email, as a few have ended up there. Looking forward to seeing you at NEMI! 🎉
021
Aaron Mueller @amuuueller.bsky.social · 27/07/2026
Tired of writing NeurIPS rebuttals? Take a break by registering to attend the New England Mechanistic Interpretability workshop (deadline today)!
010
Reposted by Aaron Mueller
New England Mechanistic Interpretability Workshop @nemiworkshop.bsky.social · 15/07/2026
The NEMI crew are entering the boat parade at Sail Boston while we wait for the mech interp abstracts to sail in on August 1
022
Reposted by Aaron Mueller
New England Mechanistic Interpretability Workshop @nemiworkshop.bsky.social · 14/07/2026
NEMI is exactly one month away! Get your abstracts in by August 1st
042
Aaron Mueller @amuuueller.bsky.social · 29/06/2026
If you'll be at ACL or ICML this year, come check out the work from our group and collaborators - summary 🧵 below. Lots to like for those into {mechanistic, developmental, pragmatic} interpretability! I'll be at ACL; say hi!
280
Aaron Mueller @amuuueller.bsky.social · 10/06/2026
The New England Mechanistic Interpretability (NEMI) workshop is coming to BU on Aug. 14! Join us for talks, a panel, food, and plenty of opportunities to connect with the many great researchers in the area. Register and help spread the word!
0176
Reposted by Aaron Mueller
Naomi Saphra @nsaphra.bsky.social · 09/06/2026
✨ it's coming ✨ NEMI 2026 will be lit. It will also be the new BU interp supergroup's debut ball. Come meet us!
nemiconf.github.io
The 3rd New England Mechanistic Interpretability (NEMI) Workshop
1274
Reposted by Aaron Mueller
Computational Linguistics Journal @complingjournal.bsky.social · 13/05/2026
Interpretability provides a toolset for understanding how and why LMs behave in certain ways. This survey proposes a perspective on interpretability research grounded in causal mediation analysis: doi.org/10.1162/COLI... #NLProc #CLJournal @jannikbrinkmann.bsky.social @amuuueller.bsky.social
0111
Reposted by Aaron Mueller
Micah Benson @micahben.bsky.social · 25/03/2026
I truly believe the rapid advances in the mech interp subfield have something real to offer AI ethics researchers: A chance to look beyond the HOW of evals to the WHY, a first pass at a technical solution when we see the opportunity, a new avenue for showing failures that prove models are not gods
173
Aaron Mueller @amuuueller.bsky.social · 21/01/2026
Representation steering is now a common way to mitigate LLM shortcuts. How much legitimate knowledge does this tend to remove? Turns out that these methods can be surprisingly precise! But also: no single steering operation will fix all shortcuts. Led by @shanzzyy.bsky.social!
060
Reposted by Aaron Mueller
languagemit.bsky.social @languagemit.bsky.social · 24/12/2025
New book! I have written a book, called Syntax: A cognitive approach, published by MIT Press. This is open access; MIT Press will post a link soon, but until then, the book is available on my website: tedlab.mit.edu/tedlab_websi...
tedlab.mit.edu
212541
Reposted by Aaron Mueller
Najoung Kim @najoung.bsky.social · 19/11/2025
I also want to mention that the lang x computation research community at BU is growing in an exciting direction, especially with new faculty like @amuuueller.bsky.social, @anthonyyacovone.bsky.social, @nsaphra.bsky.social, & @profsophie.bsky.social! Also, Boston is quite nice :)
192
Aaron Mueller @amuuueller.bsky.social · 14/11/2025
In LLMs, concepts aren’t static: they evolve through time and have rich temporal dependencies. We introduce Temporal Feature Analysis (TFA) to separate what's inferred from context vs. novel information. A big effort led by @ekdeepl.bsky.social, @sumedh-hindupur.bsky.social, @canrager.bsky.social!
1214
Reposted by Aaron Mueller
INTERPLAY Workshop@COLM '25 @interplay-workshop.bsky.social · 09/10/2025
✨ The schedule for our INTERPLAY workshop at COLM is live! ✨ 🗓️ October 10th, Room 518C 🔹 Invited talks from @sarah-nlp.bsky.social John Hewitt @amuuueller.bsky.social @kmahowald.bsky.social 🔹 Paper presentations and posters 🔹 Closing roundtable discussion. Join us in Montréal! @colmweb.org
Schedule for the INTERPLAY workshop at COLM on October 10th, Room 518C.

09:00 am: Opening
09:10 am: Invited Talks by Sarah Wiegreffe and John Hewitt
10:20 am: Paper Presentations

Lunch Break

01:00 pm: Invited Talks by Aaron Mueller and Kyle Mowhald
02:10 pm: Poster Session
03:20 pm: Roundtable Discussion
04:50 pm: Closing
044
Aaron Mueller @amuuueller.bsky.social · 01/10/2025
What's the right unit of analysis for understanding LLM internals? We explore in our mech interp survey (a major update from our 2024 ms). We’ve added more recent work and more immediately actionable directions for future work. Now published in Computational Linguistics!
24115
Reposted by Aaron Mueller
Andrew Lampinen @lampinen.bsky.social · 05/08/2025
In neuroscience, we often try to understand systems by analyzing their representations — using tools like regression or RSA. But are these analyses biased towards discovering a subset of what a system represents? If you're interested in this question, check out our new commentary! Thread:
What do representations tell us about a system? Image of a mouse with a scope showing a vector of activity patterns, and a neural network with a vector of unit activity patterns
Common analyses of neural representations: Encoding models (relating activity to task features) drawing of an arrow from a trace saying [on_____on____] to a neuron and spike train. Comparing models via neural predictivity: comparing two neural networks by their R^2 to mouse brain activity. RSA: assessing brain-brain or model-brain correspondence using representational dissimilarity matrices
617453
Aaron Mueller @amuuueller.bsky.social · 17/07/2025
If you're at #ICML2025, chat with me, @sarah-nlp.bsky.social, Atticus, and others at our poster 11am - 1:30pm at East #1205! We're establishing a 𝗠echanistic 𝗜nterpretability 𝗕enchmark. We're planning to keep this a living benchmark; come by and share your ideas/hot takes!
0133
Reposted by Aaron Mueller
David Bau @davidbau.bsky.social · 25/06/2025
The new "Lookback" paper from @nikhil07prakash.bsky.social‬ contains a surprising insight... 70b/405b LLMs use double pointers, akin to C programmers' double (**) pointers. They show up when the LLM is "knowing what Sally knows Ann knows", i.e., Theory of Mind. bsky.app/profile/nik...
bsky.app
@nikhil07prakash.bsky.social
How do language models track mental states of each character in a story, often referred to as Theory of Mind? We reverse-engineered how LLaMA-3-70B-Instruct handles a belief-tracking task and found something surprising: it uses mechanisms strikingly similar to pointer variables in C programming!
1293
Aaron Mueller @amuuueller.bsky.social · 27/05/2025
SAEs have been found to massively underperform supervised methods for steering neural networks. In new work led by @danaarad.bsky.social, we find that this problem largely disappears if you select the right features!
1161
Reposted by Aaron Mueller
Dana Arad @danaarad.bsky.social · 27/05/2025
Tried steering with SAEs and found that not all features behave as expected? Check out our new preprint - "SAEs Are Good for Steering - If You Select the Right Features" 🧵
2186
Reposted by Aaron Mueller
Ethan Gotlieb Wilcox @wegotlieb.bsky.social · 12/05/2025
Couldn’t be happier to have co-authored this will a stellar team, including: Michael Hu, @amuuueller.bsky.social, @alexwarstadt.bsky.social, @lchoshen.bsky.social, Chengxu Zhuang, @adinawilliams.bsky.social, Ryan Cotterell, @tallinzen.bsky.social
131
Aaron Mueller @amuuueller.bsky.social · 23/04/2025
Lots of progress in mech interp (MI) lately! But how can we measure when new mech interp methods yield real improvements over prior work? We propose 😎 𝗠𝗜𝗕: a 𝗠echanistic 𝗜nterpretability 𝗕enchmark!
Logo for MIB: A Mechanistic Interpretability Benchmark
15115
Aaron Mueller @amuuueller.bsky.social · 11/03/2025
Lots of work coming soon to @iclr-conf.bsky.social and @naaclmeeting.bsky.social in April/May! Come chat with us about new methods for interpreting and editing LLMs, multilingual concept representations, sentence processing mechanisms, and arithmetic reasoning. 🧵
1206
Aaron Mueller @amuuueller.bsky.social · 06/03/2025
It’s common to assume one task = one static circuit. But this ignores that the important computations depend on position! We propose a way to find 𝗽𝗼𝘀𝗶𝘁𝗶𝗼𝗻-𝗮𝘄𝗮𝗿𝗲 circuits. (Highlight: using LLMs to help us create multi-token causal abstractions!)
050
Aaron Mueller @amuuueller.bsky.social · 19/12/2024
What can mechanistic interpretability do for computational psycholinguists? @michaelwhanna.bsky.social and I took a stab at this question! We investigate garden path sentence processing in LMs at the feature (circuit) level.
1110
Reposted by Aaron Mueller
GroNLP @gronlp.bsky.social · 22/11/2024
Today we had the pleasure to host @michaelwhanna.bsky.social from @amsterdamnlp.bsky.social for a seminar on his work w/ @amuuueller.bsky.social to interpret incremental sentence processing in language models. Thank you for joining us Michael!
1224
Reposted by Aaron Mueller
Leshem (Legend) Choshen @EMNLP @lchoshen.bsky.social · 19/11/2024
The babies are now in their beds, but what a year it was! Highlighting some findings of BabyLM Architectures & Training objective matter a lot (and got the highest scores) alphaxiv.org/pdf/2410.24159 🤖
2292
Reposted by Aaron Mueller
Juan Diego Rodriguez @juand-r.bsky.social · 08/11/2024
How do language models organize concepts and their properties? Do they use taxonomies to infer new properties, or infer based on concept similarities? Apparently, both! 🌟 New paper with my fantastic collaborators @amuuueller.bsky.social and @kanishka.bsky.social
Title: "Characterizing the Role of Similarity in the Property Inferences of Language Models"
Authors: Juan Diego Rodriguez, Aaron Mueller, Kanishka Misra

Left figure: "Given that dogs are daxable, is it true that corgis are daxable?" A language model could answer this either using taxonomic relations, illustrated by a taxonomy dog-corgi, dog-mutt, canine-wolf, etc., or by similarity relations (dogs are more similar to corgis than cats, wolves or shar peis).

Right figure: illustration of the causal model (and an example intervention) for distributed alignment search (DAS), which we used to find a subspace in the network responsible for property inheritance behavior. The bottom nodes are "property", "premise concept (A)" and "conclusion concept (B)" , the middle nodes are "A has property P", "B is a kind of A", and the top node is "B has property P".
410822