Sign in

Andrew Lee

@ajyl.bsky.social
751 followers 599 following 45 posts

Post-doc @ Harvard. PhD UMich. Spent time at FAIR and MSR. ML/NLP/Interpretability

PostsRepliesMedia
Andrew Lee @ajyl.bsky.social · 09/04/2026
If you liked Anthropic's recent emotions paper, check out our work! We find many similarities: 1) Circular geometry of emotion representations 2) Steering: unlike Anthropic, we steer along circular manifold (at 0°, 30°, 60°...) 3) Steering emotions can affect refusal/sycophancy See Lihao's thread!👇
0102
Andrew Lee @ajyl.bsky.social · 30/03/2026
CFP: mechinterpworkshop.com Contact: mechinterpworkshop@gmail.com Deadline: May 8 (FYI: we are accepting Neurips format submissions 😉) Looking forward to see everyone’s work! –@@neelnanda.bsky.social, Andy Arditi, Stefan Heimersheim, Anna Soligo, @jrosseruk.bsky.social, Iván Arcuschin
mechinterpworkshop.com
Mechanistic Interpretability Workshop at ICML 2026
The Mechanistic Interpretability Workshop at ICML 2026. How can we use the internals of neural networks to understand a model better?
010
Andrew Lee @ajyl.bsky.social · 30/03/2026
We are thrilled to host the next Mech Interp Workshop @ ICML 2026! 🎉 July 2026, Seoul 🇰🇷 The workshop aims to understand the inner workings of neural nets. Topics: Feature geometry Circuit analyses Interp for {practical applications, safety, scientific discovery}, and many more.
110
Andrew Lee @ajyl.bsky.social · 09/10/2025
Thank you Naomi!!
010
Andrew Lee @ajyl.bsky.social · 04/07/2025
Question @neuripsconf.bsky.social - a coauthor had his reviews re-assigned many weeks ago. The ACs of those papers told him "i've been told to tell u: leave a short note. You won't be penalized". Now I'm being warned of desk-reject due to his short/poor reviews. What's the right protocol here?
000
Reposted by Andrew Lee
nikhil07prakash.bsky.social @nikhil07prakash.bsky.social · 24/06/2025
How do language models track mental states of each character in a story, often referred to as Theory of Mind? We reverse-engineered how LLaMA-3-70B-Instruct handles a belief-tracking task and found something surprising: it uses mechanisms strikingly similar to pointer variables in C programming!
25918
Reposted by Andrew Lee
Lihao Sun @1e0sun.bsky.social · 10/06/2025
🚨New #ACL2025 paper! Today’s “safe” language models can look unbiased—but alignment can actually make them more biased implicitly by reducing their sensitivity to race-related associations. 🧵Find out more below!
1122
Andrew Lee @ajyl.bsky.social · 13/05/2025
This project was done via Arbor! arborproject.github.io Check us out to see on-going work to interp reasoning models. Thank you collaborators! Lihao Sun, @wendlerc.bsky.social , @viegas.bsky.social , @wattenberg.bsky.social Paper link: arxiv.org/abs/2504.14379 9/n
arborproject.github.io
ARBOR
010
Andrew Lee @ajyl.bsky.social · 13/05/2025
Our interpretation: ✅we find subspace critical for self-verif. ✅in our setup, prev-token heads take resid-stream into this subspace. In a different task, a diff. mechanism may be used. ✅ this subspace activates verif-related MLP weights, promoting tokens like “success” 8/n
110
Andrew Lee @ajyl.bsky.social · 13/05/2025
We find similar verif. subspaces in our base model and general reasoning model (DeepSeek R1-14B). Here we provide CountDown as a ICL task. Interestingly, in R1-14B, our interventions lead to partial success - the LM fails self-verification but then self-corrects itself. 7/n
100
Andrew Lee @ajyl.bsky.social · 13/05/2025
Our analyses meet in the middle: We use “interlayer communication channels” to rank how much each head (OV circuit) aligns with the “receptive fields” of verification-related MLP weights. Disable *three* heads → disables self-verif. and deactivates verif.-MLP weights. 6/n
100
Andrew Lee @ajyl.bsky.social · 13/05/2025
Bottom-up, we find previous-token heads (i.e., parts of induction heads) are responsible for self-verification in our setup. Disabling previous-token heads disables self-verification. 5/n
100
Andrew Lee @ajyl.bsky.social · 13/05/2025
More importantly, we can use the probe to find MLP weights related to verification. Simply check for MLP weights with high cosine similarity to our probe. Interestingly, we often see Eng. tokens for "valid direction" and Chinese tokens for "invalid direction". 4/n
210
Andrew Lee @ajyl.bsky.social · 13/05/2025
We do a “top-down” and “bottom-up” analysis. Top-down, we train a probe. We can use our probe to steer the model and trick it to have found a solution. 3/n
110
Andrew Lee @ajyl.bsky.social · 13/05/2025
CoT is unfaithful. Can we monitor inner computations in latent space instead? Case study: Let’s study self-verification! Setup: We train Qwen-3B on CountDown until mode collapse, resulting in nicely structured CoT that’s easy to parse+analyze 2/n
100
Andrew Lee @ajyl.bsky.social · 13/05/2025
🚨New preprint! How do reasoning models verify their own CoT? We reverse-engineer LMs and find critical components and subspaces needed for self-verification! 1/n
1173
Andrew Lee @ajyl.bsky.social · 09/05/2025
media.tenor.com
a man with glasses is sitting on a couch and saying pretty pretty pretty good .
ALT: a man with glasses is sitting on a couch and saying pretty pretty pretty good .
000
Andrew Lee @ajyl.bsky.social · 07/05/2025
Interesting, I didn't know that! BTW, we find similar trends in GPT2 and Gemma2
010
Andrew Lee @ajyl.bsky.social · 07/05/2025
Next time a reviewer asks, “Why didn’t you include [insert newest LM]?”, depending on your claims you could argue that your findings will generalize to other models, based on our work! Paper link: arxiv.org/abs/2503.21073
arxiv.org
Shared Global and Local Geometry of Language Model Embeddings
Researchers have recently suggested that models share common representations. In our work, we find that token embeddings of language models exhibit common geometric structure. First, we find ``global'...
1215
Andrew Lee @ajyl.bsky.social · 07/05/2025
We call this simple approach Emb2Emb: Here we steer Llama8B using steering vectors from Llama1B and 3B:
120
Andrew Lee @ajyl.bsky.social · 07/05/2025
Now, steering vectors can be transferred across LMs. Given LM1, LM2 & their embeddings E1, E2, fit a linear transform T from E1 to E2. Given steering vector V for LM1, apply T onto V, and now TV can steer LM2. Unembedding V or TV shows similar nearest neighbors encoding the steer vector’s concept:
120
Andrew Lee @ajyl.bsky.social · 07/05/2025
Local2: We measure intrinsic dimension (ID) of token embeddings. Interestingly, ID reveals that tokens with low ID form very coherent semantic clusters, while tokens with higher ID do not!
131
Andrew Lee @ajyl.bsky.social · 07/05/2025
Local: we characterize two ways: first using Locally Linear Embeddings (LLE), in which we express each token embedding as the weighted sum of its k-nearest neighbors. It turns out, the LLE weights for most tokens look very similar across language models, indicating similar local geometry:
140
Andrew Lee @ajyl.bsky.social · 07/05/2025
We characterize “global” and “local” geometry in simple terms. Global: how similar are the distance matrices of embeddings across LMs? We can check with Pearson correlation between distance matrices: high correlation indicates similar relative orientations of token embeddings, which is what we find
141
Andrew Lee @ajyl.bsky.social · 07/05/2025
🚨New Preprint! Did you know that steering vectors from one LM can be transferred and re-used in another LM? We argue this is because token embeddings across LMs share many “global” and “local” geometric similarities!
36113
Andrew Lee @ajyl.bsky.social · 31/03/2025
Cool! QQ: say I have a "mech-interpy finding": for instance, say I found a "circuit" - is such a finding appropriate to submit, or is the workshop exclusively looking for actionable insights?
020
Andrew Lee @ajyl.bsky.social · 23/03/2025
I think these papers are highly relevant but missing! aclanthology.org/2023.emnlp-m... arxiv.org/pdf/2403.07687
aclanthology.org
Bridging the Digital Divide: Performance Variation across Socio-Economic Factors in Vision-Language Models
Joan Nwatu, Oana Ignat, Rada Mihalcea. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
120
Andrew Lee @ajyl.bsky.social · 05/03/2025
My website has a personal readme file with step by step instructions on how to make an update. I would need to hire someone if something were to ever happen to that readme file.
010
Reposted by Andrew Lee
David Bau @davidbau.bsky.social · 20/02/2025
Today we launch a new open research community It is called ARBOR: arborproject.github.io/ please join us. bsky.app/profile/ajy...
1155
Andrew Lee @ajyl.bsky.social · 20/02/2025
Check out on-going projects here! github.com/ARBORproject... or join our discord: discord.gg/SeBdQbRPkA We hope to see your contributions! Organizers: @wattenberg.bsky.social @viegas.bsky.social @davidbau.bsky.social @wendlerc.bsky.social @canrager.bsky.social 7/N
github.com
ARBORproject arborproject.github.io · Discussions
Explore the GitHub Discussions forum for ARBORproject arborproject.github.io. Discuss code, ask questions & collaborate with the developer community.
020
Andrew Lee @ajyl.bsky.social · 20/02/2025
@canrager.bsky.social finds that R1’s reasoning tokens reveal a lot about restricted topics. Can we find all restricted topics of a reasoning model, by understanding its inner mechanisms? dsthoughts.baulab.info 6/N
dsthoughts.baulab.info
Auditing AI Bias: The DeepSeek Case
Cracking open the inner monologue of reasoning models.
100
Andrew Lee @ajyl.bsky.social · 20/02/2025
Similarly, I find linear vectors that seem to encode whether R1 has found a solution or not. We can steer the model to think that it’s found a solution during its CoT with simple linear interventions: ajyl.github.io/2025/02/16/s... 5/N
100
Andrew Lee @ajyl.bsky.social · 20/02/2025
We have some preliminary findings. @wendlerc.bsky.social find simple steering vectors to either make the model continue its CoT (ie, double check its answers) or finish its thought on a GSM8K github.com/ARBORproject...
100
Andrew Lee @ajyl.bsky.social · 20/02/2025
Contributions can include: experiments for an on-going project; new resources (model/SAE checkpoints) datasets, or even software (general infra, tools for data collection/annotation, etc.). Of course feel free to launch your own projects! See on-going projects: github.com/ARBORproject... 3/N
github.com
ARBORproject arborproject.github.io · Discussions
Explore the GitHub Discussions forum for ARBORproject arborproject.github.io. Discuss code, ask questions & collaborate with the developer community.
100
Andrew Lee @ajyl.bsky.social · 20/02/2025
We draw inspiration from previous open-research (eg teorth.github.io/equational_t...) to accelerate progress. Our goal is to facilitate connections between researchers, minimize duplicate efforts, and host a centralized repository of resources. 2/N
teorth.github.io
Equational Theories Project
Mapping out the relations between different equational theories of Magmas
100
Andrew Lee @ajyl.bsky.social · 20/02/2025
Excited about recent reasoning models? What is happening under the hood? Join ARBOR: Analysis of Reasoning Behaviors thru *Open Research* - a radically open collaboration to reverse-engineer reasoning models! Learn more: arborproject.github.io 1/N
arborproject.github.io
ARBOR
1133
Andrew Lee @ajyl.bsky.social · 05/01/2025
Hope you enjoy our work! Thank you team members for an amazing collaboration! @corefpark.bsky.social @ekdeepl.bsky.social YongyiYang MayaOkawa KentoNishi @wattenberg.bsky.social @hidenori8tanaka.bsky.social Special thanks: @ndif-team.bsky.social 9/N
010
Andrew Lee @ajyl.bsky.social · 05/01/2025
Interestingly, similar observations have been made in humans: human brains also construct graphical (spatial+relational) representations given enough observations of random images from a graph! doi.org/10.7554/eLif... 8/N
120
Andrew Lee @ajyl.bsky.social · 05/01/2025
Why does this happen? We hypothesize LMs internally perform a graph spectral energy minimization process: Thus we quantify + measure energy of model activations: when energy is reduced, in-context task accuracy *jumps* to near 100%. Phase-transition of in-context reps? 7/N
120
Andrew Lee @ajyl.bsky.social · 05/01/2025
In-context representations form across models and various graphs. We can intervene on the principal components to alter the model’s belief-state of the graph, and change the model’s in-context task predictions accordingly (ie, causality). 6/N
110
Andrew Lee @ajyl.bsky.social · 05/01/2025
Instead of random words, we also try words w/ a semantic prior (ex: days of week). We make an in-context task with a new ordering of days (Mon,Thurs,Sun,Wed…): The orig. semantic prior shows up in first principal components & in-context reps in latter ones! The two co-exist. 5/N
110
Andrew Lee @ajyl.bsky.social · 05/01/2025
@JoshAEngels find that words with a semantic prior (ex: days of the week) can form a ring (x.com/JoshAEngels/...) These rings can also be formed in-context! 4/N
120
Andrew Lee @ajyl.bsky.social · 05/01/2025
What are in-context representations? We build a graph with random words as nodes, randomly walk the graph, and input the resulting seq to LMs. With a long enough context, LMs start adhering to the graph When this happens, PCA of LM activations reveals the graph (task) structure 3/N
110
Andrew Lee @ajyl.bsky.social · 05/01/2025
Given an in-context task, and as context is scaled, LMs can form ‘in-context representations’ that reflect the task. Team: @corefpark.bsky.social @ekdeepl.bsky.social YongyiYang MayaOkawa KentoNishi @wattenberg.bsky.social @hidenori8tanaka.bsky.social arxiv.org/pdf/2501.00070 2/N
131
Andrew Lee @ajyl.bsky.social · 05/01/2025
New paper <3 Interested in inference-time scaling? In-context Learning? Mech Interp? LMs can solve novel in-context tasks, with sufficient examples (longer contexts). Why? Bc they dynamically form *in-context representations*! 1/N
25316
Andrew Lee @ajyl.bsky.social · 23/12/2024
Such social phenomenon has been observed before! journals.sagepub.com/doi/abs/10.1...
journals.sagepub.com
When Knowledge Work and Analytical Technologies Collide: The Practices and Consequences of Black Boxing Algorithmic Technologies - Callen Anthony, 2021
Analytical technologies that structure and process data hold great promise for organizations but also may pose fundamental challenges for how knowledge workers ...
260
Andrew Lee @ajyl.bsky.social · 03/12/2024
These papers seem relevant, especially the first one? arxiv.org/abs/2404.00859 and arxiv.org/abs/2311.04897
arxiv.org
Do language models plan ahead for future tokens?
Do transformers "think ahead" during inference at a given position? It is known transformers prepare information in the hidden states of the forward pass at time step $t$ that is then used in future f...
120
Andrew Lee @ajyl.bsky.social · 26/11/2024
Lol. But actually - could you point me to your favorite ones?
130