Sign in

Andrew Lee

@ajyl.bsky.social
750 followers 599 following 45 posts

Post-doc @ Harvard. PhD UMich. Spent time at FAIR and MSR. ML/NLP/Interpretability

PostsRepliesMedia
Andrew Lee @ajyl.bsky.social · 09/04/2026
If you liked Anthropic's recent emotions paper, check out our work! We find many similarities: 1) Circular geometry of emotion representations 2) Steering: unlike Anthropic, we steer along circular manifold (at 0°, 30°, 60°...) 3) Steering emotions can affect refusal/sycophancy See Lihao's thread!👇
0102
Andrew Lee @ajyl.bsky.social · 30/03/2026
We are thrilled to host the next Mech Interp Workshop @ ICML 2026! 🎉 July 2026, Seoul 🇰🇷 The workshop aims to understand the inner workings of neural nets. Topics: Feature geometry Circuit analyses Interp for {practical applications, safety, scientific discovery}, and many more.
110
Andrew Lee @ajyl.bsky.social · 04/07/2025
Question @neuripsconf.bsky.social - a coauthor had his reviews re-assigned many weeks ago. The ACs of those papers told him "i've been told to tell u: leave a short note. You won't be penalized". Now I'm being warned of desk-reject due to his short/poor reviews. What's the right protocol here?
000
Reposted by Andrew Lee
nikhil07prakash.bsky.social @nikhil07prakash.bsky.social · 24/06/2025
How do language models track mental states of each character in a story, often referred to as Theory of Mind? We reverse-engineered how LLaMA-3-70B-Instruct handles a belief-tracking task and found something surprising: it uses mechanisms strikingly similar to pointer variables in C programming!
25918
Reposted by Andrew Lee
Lihao Sun @1e0sun.bsky.social · 10/06/2025
🚨New #ACL2025 paper! Today’s “safe” language models can look unbiased—but alignment can actually make them more biased implicitly by reducing their sensitivity to race-related associations. 🧵Find out more below!
1122
Andrew Lee @ajyl.bsky.social · 13/05/2025
🚨New preprint! How do reasoning models verify their own CoT? We reverse-engineer LMs and find critical components and subspaces needed for self-verification! 1/n
1173
Andrew Lee @ajyl.bsky.social · 07/05/2025
🚨New Preprint! Did you know that steering vectors from one LM can be transferred and re-used in another LM? We argue this is because token embeddings across LMs share many “global” and “local” geometric similarities!
36113
Reposted by Andrew Lee
David Bau @davidbau.bsky.social · 20/02/2025
Today we launch a new open research community It is called ARBOR: arborproject.github.io/ please join us. bsky.app/profile/ajy...
1155
Andrew Lee @ajyl.bsky.social · 20/02/2025
Excited about recent reasoning models? What is happening under the hood? Join ARBOR: Analysis of Reasoning Behaviors thru *Open Research* - a radically open collaboration to reverse-engineer reasoning models! Learn more: arborproject.github.io 1/N
arborproject.github.io
ARBOR
1133
Andrew Lee @ajyl.bsky.social · 05/01/2025
New paper <3 Interested in inference-time scaling? In-context Learning? Mech Interp? LMs can solve novel in-context tasks, with sufficient examples (longer contexts). Why? Bc they dynamically form *in-context representations*! 1/N
25316