Sign in

Eric Todd

@ericwtodd.bsky.social
508 followers 184 following 16 posts

CS PhD Student, Northeastern University - Machine Learning, Interpretability ericwtodd.github.io

PostsRepliesMedia
Eric Todd @ericwtodd.bsky.social · 18/04/2026
I'll be attending #ICLR2026 next week to present my work on In-Context Algebra! My poster will be on Fri, April 24 at 3:15-5:45PM at Pavilion 4 P4-#4011. If you're around, stop by and say hello! My DMs are open if you want to connect or meet up in Rio! bsky.app/profile/eri...
bsky.app
Eric Todd (@ericwtodd.bsky.social)
Can you solve this algebra puzzle? 🧩 cb=c, ac=b, ab=? A small transformer can learn to solve problems like this! And since the letters don't have inherent meaning, this lets us study how context alone imparts meaning. Here's what we found:🧵⬇️
0120
Reposted by Eric Todd
Hadas Orgad @hadasorgad.bsky.social · 13/04/2026
New paper: LLMs encode harmful content generation in a distinct, unified mechanism Using weight pruning, we find that harmful generation depends on a tiny subset of the weights that are shared across harm types and separate from benign capabilities. 🧵
172
Reposted by Eric Todd
David Bau @davidbau.bsky.social · 27/01/2026
The Art of Wanting. About the question I see as central in AI ethics, interpretability, and safety. Can an AI take responsibility? I do not think so, but *not* because it's not smart enough. davidbau.com/archives/20...
193
Reposted by Eric Todd
Koyena Pal @koyena.bsky.social · 22/01/2026
Can models understand each other's reasoning? 🤔 When Model A explains its Chain-of-Thought (CoT) , do Models B, C, and D interpret it the same way? Our new preprint with @davidbau.bsky.social and @csinva.bsky.social explores CoT generalizability 🧵👇 (1/7)
1278
Eric Todd @ericwtodd.bsky.social · 22/01/2026
Can you solve this algebra puzzle? 🧩 cb=c, ac=b, ab=? A small transformer can learn to solve problems like this! And since the letters don't have inherent meaning, this lets us study how context alone imparts meaning. Here's what we found:🧵⬇️
24811
Reposted by Eric Todd
Can @canrager.bsky.social · 13/11/2025
Humans and LLMs think fast and slow. Do SAEs recover slow concepts in LLMs? Not really. Our Temporal Feature Analyzer discovers contextual features in LLMs, that detect event boundaries, parse complex grammar, and represent ICL patterns.
1208
Reposted by Eric Todd
Hiba Ahsan @hibaahsan.bsky.social · 05/11/2025
LLMs have been shown to provide different predictions in clinical tasks when patient race is altered. Can SAEs spot this undue reliance on race? 🧵 Work w/ @byron.bsky.social Link: arxiv.org/abs/2511.00177
152
Reposted by Eric Todd
Jennifer Hu @jennhu.bsky.social · 04/11/2025
Interested in doing a PhD at the intersection of human and machine cognition? ✨ I'm recruiting students for Fall 2026! ✨ Topics of interest include pragmatics, metacognition, reasoning, & interpretability (in humans and AI). Check out JHU's mentoring program (due 11/15) for help with your SoP 👇
02815
Reposted by Eric Todd
Arnab Sen Sharma @arnabsensharma.bsky.social · 04/11/2025
How can a language model find the veggies in a menu? New pre-print where we investigate the internal mechanisms of LLMs when filtering on a list of options. Spoiler: turns out LLMs use strategies surprisingly similar to functional programming (think "filter" from python)! 🧵
1249
Eric Todd @ericwtodd.bsky.social · 06/10/2025
Looking forward to attending #COLM2025 this week! Would love to meet up and chat with others about interpretability + more. DMs are open if you want to connect. Be sure to checkout @sfeucht.bsky.social's very cool work on understanding concepts in LLMs tomorrow morning (Poster 35)!
020
Reposted by Eric Todd
Aaron Mueller @amuuueller.bsky.social · 01/10/2025
What's the right unit of analysis for understanding LLM internals? We explore in our mech interp survey (a major update from our 2024 ms). We’ve added more recent work and more immediately actionable directions for future work. Now published in Computational Linguistics!
24115
Reposted by Eric Todd
David Bau @davidbau.bsky.social · 27/09/2025
Who is going to be at #COLM2025? I want to draw your attention to a COLM paper by my student @sfeucht.bsky.social that has totally changed the way I think and teach about LLM representations. The work is worth knowing. And you can meet Sheridan at COLM, Oct 7! bsky.app/profile/sfe...
1398
Reposted by Eric Todd
David Bau @davidbau.bsky.social · 26/09/2025
Announcing a broad expansion of the National Deep Inference Fabric. This could be relevant to your research...
1113
Reposted by Eric Todd
Chantal @chantalsh.bsky.social · 24/09/2025
"AI slop" seems to be everywhere, but what exactly makes text feel like "slop"? In our new work (w/ @tuhinchakr.bsky.social, Diego Garcia-Olano, @byron.bsky.social ) we provide a systematic attempt at measuring AI "slop" in text! arxiv.org/abs/2509.19163 🧵 (1/7)
13317
Reposted by Eric Todd
Millicent Li @millicentli.bsky.social · 17/09/2025
Wouldn’t it be great to have questions about LM internals answered in plain English? That’s the promise of verbalization interpretability. Unfortunately, our new paper shows that evaluating these methods is nuanced—and verbalizers might not tell us what we hope they do. 🧵👇1/8
1268
Reposted by Eric Todd
David Bau @davidbau.bsky.social · 18/08/2025
This Friday NEMI 2025 is at Northeastern in Boston, 8 talks, 24 roundtables, 90 posters; 200+ attendees. Thanks to goodfire.ai/ for sponsoring! nemiconf.github.io/summer25/ If you can't make it in person, the livestream will be here: www.youtube.com/live/4BJBis...
youtube.com
New England Mechanistic Interpretability Workshop
About:The New England Mechanistic Interpretability (NEMI) workshop aims to bring together academic and industry researchers from the New England and surround...
1167
Reposted by Eric Todd
Sheridan Feucht @sfeucht.bsky.social · 22/07/2025
We've added a quick new section to this paper, which was just accepted to @COLM_conf! By summing weights of concept induction heads, we created a "concept lens" that lets you read out semantic information in a model's hidden states. 🔎
171
Eric Todd @ericwtodd.bsky.social · 01/07/2025
Im excited for NEMI again this year! I’ve enjoyed local research meetups and getting to know others near me working on interesting problems.
010
Reposted by Eric Todd
Koyena Pal @koyena.bsky.social · 30/06/2025
🚨 Registration is live! 🚨 The New England Mechanistic Interpretability (NEMI) Workshop is happening Aug 22nd 2025 at Northeastern University! A chance for the mech interp community to nerd out on how models really work 🧠🤖 🌐 Info: nemiconf.github.io/summer25/ 📝 Register: forms.gle/v4kJCweE3UUH...
NEMI 2024 (Last Year)
0108
Reposted by Eric Todd
nikhil07prakash.bsky.social @nikhil07prakash.bsky.social · 24/06/2025
How do language models track mental states of each character in a story, often referred to as Theory of Mind? We reverse-engineered how LLaMA-3-70B-Instruct handles a belief-tracking task and found something surprising: it uses mechanisms strikingly similar to pointer variables in C programming!
25918
Reposted by Eric Todd
Can @canrager.bsky.social · 13/06/2025
Can we uncover the list of topics a language model is censored on? Refused topics vary strongly among models. Claude-3.5 vs DeepSeek-R1 refusal patterns:
1104
Reposted by Eric Todd
Sheridan Feucht @sfeucht.bsky.social · 10/04/2025
I'll present a poster for this work at NENLP tomorrow! Come find me at poster #80...
071
Reposted by Eric Todd
David Bau @davidbau.bsky.social · 07/04/2025
Sheridan asks whether the Dual Route Model of Reading that psychologists have observed in humans also appears in LLMs. In her brilliantly simple study of induction heads, she finds that it does! Induction has a Dual Route that separates concepts from literal token processing. Worth reading ↘️
072
Reposted by Eric Todd
Sheridan Feucht @sfeucht.bsky.social · 07/04/2025
[📄] Are LLMs mindless token-shifters, or do they build meaningful representations of language? We study how LLMs copy text in-context, and physically separate out two types of induction heads: token heads, which copy literal tokens, and concept heads, which copy word meanings.
17518
Reposted by Eric Todd
David Bau @davidbau.bsky.social · 16/03/2025
What will be the linchpin for AI dominance? Read our NSF/OSTP recommendations written with Goodfire's Tom McGrath tommcgrath.github.io, Transluce's Sarah Schwettmann cogconfluence.com, MIT's Dylan Hadfield-Menell @dhadfieldmenell.bsky.social TLDR; Dominance comes from **interpretability** 🧵 ↘️
1218
Reposted by Eric Todd
Chantal @chantalsh.bsky.social · 10/03/2025
I'm searching for some comp/ling experts to provide a precise definition of “slop” as it refers to text (see: corp.oup.com/word-of-the-...) I put together a google form that should take no longer than 10 minutes to complete: forms.gle/oWxsCScW3dJU... If you can help, I'd appreciate your input! 🙏
corp.oup.com
Oxford Word of the Year 2024 - Oxford University Press
The Oxford Word of the Year 2024 is 'brain rot'. Discover more about the winner, our shortlist, and 20 years of words that reflect the world.
0108
Reposted by Eric Todd
Kayo Yin @kayoyin.bsky.social · 28/02/2025
Induction heads are commonly associated with in-context learning, but are they the primary driver of ICL at scale? We find that recently discovered "function vector" heads, which encode the ICL task, are the actual primary mechanisms behind few-shot ICL! arxiv.org/abs/2502.14010 🧵👇
1227
Reposted by Eric Todd
Hiba Ahsan @hibaahsan.bsky.social · 22/02/2025
LLMs are known to perpetuate social biases in clinical tasks. Can we locate and intervene upon LLM activations that encode patient demographics like gender and race? 🧵 Work w/ @arnabsensharma.bsky.social, @silvioamir.bsky.social, @davidbau.bsky.social, @byron.bsky.social arxiv.org/abs/2502.13319
3177
Reposted by Eric Todd
NDIF Team @ndif-team.bsky.social · 20/02/2025
Please help amplify ARBOR, a fantastic new research opportunity! If you’d like to start contributing, NDIF is now hosting DeepSeek R1 8B and 70B, open for all researchers to experiment on via our API. Sign up for API access here: login.ndif.us
login.ndif.us
NDIF Registration
043
Eric Todd @ericwtodd.bsky.social · 20/02/2025
I'm excited about this new open research initiative! It kind of feels like this is how science is supposed to be done - collaborating and sharing ideas in the open. If you've thought about studying the mechanisms behind R1 & other reasoning models check it out!
000
Reposted by Eric Todd
David Bau @davidbau.bsky.social · 31/01/2025
DeepSeek R1 shows how important it is to be studying the internals of reasoning models. Try our code: Here @canrager.bsky.social shows a method for auditing AI bias by probing the internal monologue. dsthoughts.baulab.info I'd be interested in your thoughts.
dsthoughts.baulab
1289
Reposted by Eric Todd
David Bau @davidbau.bsky.social · 31/12/2024
What was the most important machine learning paper in 2024? My Famous Deep Learning Papers list (that I use in teaching) does not include any new ideas from the last year. papers.baulab.info Which single new paper would you add?
105511
Reposted by Eric Todd
NDIF Team @ndif-team.bsky.social · 10/12/2024
More big news! Applications are open for the NDIF Summer Engineering Fellowship—an opportunity to work on cutting-edge AI research infrastructure this summer in Boston! 🚀
196
Reposted by Eric Todd
David Bau @davidbau.bsky.social · 10/12/2024
The Phase 2 NDIF Pilot is open for a short window. Apply now to get research capacity on Llama 405b. Deadline is December 31. It is not easy to crack open 405b for research, but NDIF solves the key engineering problems for you. Phase 1 powered several very interesting ICLR submissions...
0145
Reposted by Eric Todd
NDIF Team @ndif-team.bsky.social · 09/12/2024
Do you have a great experiment that you want to run on Llama 405b but not enough GPUs? 🚨 #NDIF is opening up more spots in our 405b pilot program! Apply now for a chance to conduct your own groundbreaking experiments on the 405b model. Details: 🧵⬇️
1184
Reposted by Eric Todd
David Bau @davidbau.bsky.social · 07/12/2024
PhD Applicants: remember that the Northeastern Computer Science PhD application deadline is Dec 15. It's a terrific time to do a PhD, with so many interesting things happening in AI. Apply here: www.khoury.northeastern.edu/apply/phd-ap...
khoury.northeastern.edu
PhD Apply - Khoury College of Computer Sciences
0335
Reposted by Eric Todd
Rohit Gandikota @rohitgandikota.bsky.social · 04/12/2024
New Preprint 🚀 Can diffusion models draw artistic inspiration from nature? 🤔 @huiren and @materzynska trained a diffusion model solely on natural images, excluding artwork from pre-training data ❌🎨 Surprisingly, it can mimic art styles! Curious how it works?👇🧵w/ @davidbau and Antonio Torralba
131
Reposted by Eric Todd
David Bau @davidbau.bsky.social · 04/12/2024
Do you need to copy art to make art? Hui Ren's and Joanna Materzynska's Art-Free Diffusion tests this question and lets you make "imitation-free" AI art Github: github.com/rhfeiyang/ar... Arxiv: arxiv.org/abs/2412.00176 Website: joaanna.github.io/art-free-dif... X: x.com/materzynska/...
github.com
GitHub - rhfeiyang/art-free-diffusion: Official implementation of "Art-Free Generative Models: Art Creation Without Graphic Art Knowledge"
Official implementation of "Art-Free Generative Models: Art Creation Without Graphic Art Knowledge" - rhfeiyang/art-free-diffusion
0172