Sign in

Sheridan Feucht

@sfeucht.bsky.social
255 followers 334 following 40 posts

PhD student doing LLM interpretability with @davidbau.bsky.social and @byron.bsky.social. (they/them) sfeucht.github.io

PostsRepliesMedia
Sheridan Feucht @sfeucht.bsky.social · 28/09/2026
Thanks to co-authors @bennokrojer.bsky.social, Sarah Wang, Henry Abrahamsen, @byron.bsky.social, and @davidbau.bsky.social for helping get this paper done. :) Website: ocr.baulab.info/ arXiv: arxiv.org/abs/2609.18823 Code: github.com/sfeucht/ocr
github.com
GitHub - sfeucht/ocr: Code for https://ocr.baulab.info/
Code for https://ocr.baulab.info/. Contribute to sfeucht/ocr development by creating an account on GitHub.
040
Sheridan Feucht @sfeucht.bsky.social · 28/09/2026
If you want to try the lens for yourself with a few pre-computed example images across four models, check out our demo on the paper website! ocr.baulab.info/#demo
100
Sheridan Feucht @sfeucht.bsky.social · 28/09/2026
Here's a full edit example: wherever Qwen latently represents "tractor," we replace it with the "revolver" direction. Since the model can still see features like wheels and a license plate, it interprets the image not as a tractor, but as a "surreal revolver that looks like a car"!
100
Sheridan Feucht @sfeucht.bsky.social · 28/09/2026
Is the info being read by this lens actually causal, though? To answer this, we use the *inverse* of our lens to obtain latent concept vectors. Just by adding the latent vector for "iPod" and subtracting the vector for "ant," we can fool Qwen into describing an ant as an iPod.
110
Sheridan Feucht @sfeucht.bsky.social · 28/09/2026
Some of the information revealed by this lens is local (e.g., colors, objects), but other information seems to be broader. For example, Qwen3-8B seems to think of this photo as being taken in Wyoming (which isn't too far off from where I actually took the photo, in Calgary).
100
Sheridan Feucht @sfeucht.bsky.social · 28/09/2026
This method gives us probs. across all vocab tokens, so we can observe subtleties depending on word meaning. For example, "phone" has high probability in this image for both a cell phone and landline, whereas 手机 and 电话 focus on cell phone / landline tokens respectively.
100
Sheridan Feucht @sfeucht.bsky.social · 28/09/2026
To quickly run this approach over many image tokens, we collapse together the weight matrices (the OV matrices) of this set of verbalization heads. This gives us a single d_model x d_model matrix that can be applied to any hidden state at any layer, and even works at layer 0.
100
Sheridan Feucht @sfeucht.bsky.social · 28/09/2026
First, we found some heads responsible for OCR (see paper). We could see them attending to image tokens that contain text, e.g. the word "bike". But then, we found that if we just forced all of these heads to attend to, e.g., a bird wing token, Qwen3 would output "feathers." 😮
100
Sheridan Feucht @sfeucht.bsky.social · 28/09/2026
[New preprint] How do images in VLMs align with words? In this work, we found a set of attention heads responsible for OCR. But to our surprise, these heads were actually able to verbalize much more than just text. So, we used them to create a simple logit lens for image tokens!
1107
Sheridan Feucht @sfeucht.bsky.social · 10/08/2026
Just look... you know what I mean? github.com/EliSchwartz/...
github.com
000
Sheridan Feucht @sfeucht.bsky.social · 10/08/2026
Logging back on here just to make an ImageNet appreciation post. It's so refreshing to look at *real* random images instead of diffusion-generated ones. If you prompted FLUX to generate a photo of a desktop computer, you would never see that wallpaper... real images have some special quality to them
110
Reposted by Sheridan Feucht
Koyena Pal @koyena.bsky.social · 22/01/2026
Can models understand each other's reasoning? 🤔 When Model A explains its Chain-of-Thought (CoT) , do Models B, C, and D interpret it the same way? Our new preprint with @davidbau.bsky.social and @csinva.bsky.social explores CoT generalizability 🧵👇 (1/7)
1278
Sheridan Feucht @sfeucht.bsky.social · 22/01/2026
Eric’s setting is a really cool way of studying “pure” in-context learning, where symbols always mean different things in different contexts (so the model has to constantly do ICL). He has some really cool experiments you should check out 🔎
050
Sheridan Feucht @sfeucht.bsky.social · 14/12/2025
Would love to chat with people about this, especially on the philosophy side of things--I don't have a background in philosophy, but I've been stuck on these ideas for a few years now.
000
Sheridan Feucht @sfeucht.bsky.social · 14/12/2025
I just finished writing a blog post about a fake ICLR paper that caused a stir about a month ago. I think it has interesting implications for whether AI-generated text is "meaningful" or not. sfeucht.github.io/rerereading/
120
Sheridan Feucht @sfeucht.bsky.social · 06/12/2025
I'm headed to San Diego this afternoon to attend the mech interp workshop at #NeurIPS! Really excited to give a lightning talk and catch up with everyone. 🌴
041
Reposted by Sheridan Feucht
Kanishka Misra @kanishka.bsky.social · 28/07/2025
Looking forward to attending #cogsci2025 (Jul 29 - Aug 3)! I’m especially excited to meet students who will be applying to PhD programs in Computational Ling/CogSci in the coming cycle. Please reach out if you want to meet up and chat! Email is the best way, but DM also works if you must! quick🧵:
Placeholders for 3 students (number arbitrarily chosen) and me - to signify my eventual group!
1217
Sheridan Feucht @sfeucht.bsky.social · 22/07/2025
Try it out in our new paper demo notebook! Or ping me with any sequence to try and I'd be more than happy to run a few examples for you. colab.research.google.com/github/sfeuc... Also check out the new camera-ready version of the paper on arXiv. arxiv.org/abs/2504.03022
colab.research.google.com
Google Colab
010
Sheridan Feucht @sfeucht.bsky.social · 22/07/2025
If we do the same for token induction heads, we can also get a "token lens", which reads out surface-level token information from states. Unlike raw logit lens, which reveals next-token predictions, "token lens" reveals the current token.
"Token lens" outputs for the token "card" in the context "in the morning air, she heard northern card.inals."
100
Sheridan Feucht @sfeucht.bsky.social · 22/07/2025
If we apply concept lens to the word "cardinals" in three contexts, we see that Llama-2-7b has encoded this word very differently in each case!
Three "concept lens" outputs, showing the top-5 highest probability tokens when a hidden state (throughout different layers) is transformed by concept lens and projected to token space. There are three sentences, each with different predictions: "he was a lifelong fan of the cardinals", for which concept lens predicts "football" and "baseball"; "the secret meeting of the cardinals", for which concept lens predicts "Catholic"; and "in the morning air, she hear northern cardinals", which projects to "birds."
110
Sheridan Feucht @sfeucht.bsky.social · 22/07/2025
To do this, we sum the OV matrices of the top-k concept induction heads, and use it to transform a hidden state at a particular token position. Projecting that to vocab space with the model's decoder head, we can access the "meaning" encoded in that state.
100
Sheridan Feucht @sfeucht.bsky.social · 22/07/2025
We've added a quick new section to this paper, which was just accepted to @COLM_conf! By summing weights of concept induction heads, we created a "concept lens" that lets you read out semantic information in a model's hidden states. 🔎
171
Reposted by Sheridan Feucht
Koyena Pal @koyena.bsky.social · 30/06/2025
🚨 Registration is live! 🚨 The New England Mechanistic Interpretability (NEMI) Workshop is happening Aug 22nd 2025 at Northeastern University! A chance for the mech interp community to nerd out on how models really work 🧠🤖 🌐 Info: nemiconf.github.io/summer25/ 📝 Register: forms.gle/v4kJCweE3UUH...
NEMI 2024 (Last Year)
0108
Sheridan Feucht @sfeucht.bsky.social · 24/06/2025
Nikhil's recent paper is a tour de force in causal analysis! They show that LLMs keep track of what characters know in a story using "pointer" mechanisms. Definitely worth checking out.
042
Sheridan Feucht @sfeucht.bsky.social · 25/04/2025
I'm on the train right now and just finished reading this paper for the first time--I actually just logged back on to bsky just so that I could link to it, but you beat me to the punch! I really enjoyed your paper. This example was particularly great.
010
Sheridan Feucht @sfeucht.bsky.social · 25/04/2025
I used to think formal reasoning was central to language and intelligence, but now I’m not so sure. Wrote a short post about my thoughts on this, with a couple chewy anecdotes. Would love to get some feedback or pointers to further reading. sfeucht.github.io/syllogisms/
sfeucht.github.io
Sheridan Feucht
Solving Syllogisms is Not Intelligence April 23, 2025 (I think that we overvalue logical reasoning when it comes to measuring "intelligence.") What do we mean by intelligence in the context of cogniti...
160
Sheridan Feucht @sfeucht.bsky.social · 10/04/2025
I'll present a poster for this work at NENLP tomorrow! Come find me at poster #80...
071
Sheridan Feucht @sfeucht.bsky.social · 09/04/2025
That’s a good point! Sort of related, I noticed last night that when I have to type in a 2FA code I usually compress the numbers. Like if the code is 51692 I think “fifty-one, sixty-nine, two.” I wonder if this is a thing that people have studied. Thanks for the comment :)
010
Sheridan Feucht @sfeucht.bsky.social · 07/04/2025
Paper: arxiv.org/abs/2504.03022 Code: github.com/sfeucht/dual... See dualroute.baulab.info for more info. Work done with @ericwtodd.bsky.social, @byron.bsky.social, and @davidbau.bsky.social. :)
arxiv.org
The Dual-Route Model of Induction
Prior work on in-context copying has shown the existence of induction heads, which attend to and promote individual tokens during copying. In this work we introduce a new type of induction head: conce...
020
Sheridan Feucht @sfeucht.bsky.social · 07/04/2025
Yin & Steinhardt (2025) recently showed that FV heads are more important for ICL than token induction heads. But for translation, *concept* induction heads matter too! They copy forward word meanings, whereas FV heads influence the output language. bsky.app/profile/kay...
120
Sheridan Feucht @sfeucht.bsky.social · 07/04/2025
Concept heads also output language-agnostic word representations. If we patch the outputs of these heads from one translation prompt to another, we can change the *meaning* of the outputted word, without changing the language. (see prior work from @butanium.bsky.social and @wendlerc.bsky.social)
140
Sheridan Feucht @sfeucht.bsky.social · 07/04/2025
Token induction heads are still important, though. When we ablate them over long sequences, models start to paraphrase instead of copying. We take this to mean that token induction heads are responsible for *exact* copying (which concept induction heads apparently can't do).
120
Sheridan Feucht @sfeucht.bsky.social · 07/04/2025
But how do we know these heads copy semantics? When we ablate concept induction heads, performance drops drastically for translation, synonyms, and antonyms: all tasks that require copying *meaning*, not just literal tokens.
130
Sheridan Feucht @sfeucht.bsky.social · 07/04/2025
Previous work showed that token induction heads attend to the next token to be copied (*window*pane). Analogously, we find that concept induction heads attend to the end of the next multi-token word to be copied (windowp*ane*).
120
Sheridan Feucht @sfeucht.bsky.social · 07/04/2025
--using causal interventions. Essentially, we pick out all of the attention heads that are responsible for promoting future entity tokens (e.g. "ax" in "waxwing"). We hypothesize that heads carrying an entire entity actually represent the *meaning* of that chunk of tokens.
120
Sheridan Feucht @sfeucht.bsky.social · 07/04/2025
Induction heads were discovered by Elhage et al. (2021) and Olsson et al. (2022). They focused on token copying, but some of the heads they found also seemed to activate for "fuzzy" copying tasks, like translation. We directly identify these heads-- transformer-circuits.pub/2022/in-con...
130
Sheridan Feucht @sfeucht.bsky.social · 07/04/2025
There are multiple ways to copy text! Copying a wifi password like hxioW2qN52 is different than copying a meaningful one like OwlDoorGlass. Nonsense copying requires each char to be transferred one-by-one, but meaningful words can be copied all at once. Turns out, LLMs do both.
251
Sheridan Feucht @sfeucht.bsky.social · 07/04/2025
[📄] Are LLMs mindless token-shifters, or do they build meaningful representations of language? We study how LLMs copy text in-context, and physically separate out two types of induction heads: token heads, which copy literal tokens, and concept heads, which copy word meanings.
17518
Sheridan Feucht @sfeucht.bsky.social · 02/04/2025
So gorgeous, is this in Cambridge?
110
Sheridan Feucht @sfeucht.bsky.social · 12/03/2025
Looks really cool! Can’t wait to give this a proper read.
040
Reposted by Sheridan Feucht
Chantal @chantalsh.bsky.social · 10/03/2025
I'm searching for some comp/ling experts to provide a precise definition of “slop” as it refers to text (see: corp.oup.com/word-of-the-...) I put together a google form that should take no longer than 10 minutes to complete: forms.gle/oWxsCScW3dJU... If you can help, I'd appreciate your input! 🙏
corp.oup.com
Oxford Word of the Year 2024 - Oxford University Press
The Oxford Word of the Year 2024 is 'brain rot'. Discover more about the winner, our shortlist, and 20 years of words that reflect the world.
0108
Sheridan Feucht @sfeucht.bsky.social · 22/02/2025
I like this work a lot. Racism+misogyny in medicine is genuinely dangerous, so it's really important to keep tabs on model biases if we're going to use LLMs in clinical settings. It's nice to see that interpretability techniques are useful here.
060
Reposted by Sheridan Feucht
NDIF Team @ndif-team.bsky.social · 09/12/2024
Do you have a great experiment that you want to run on Llama 405b but not enough GPUs? 🚨 #NDIF is opening up more spots in our 405b pilot program! Apply now for a chance to conduct your own groundbreaking experiments on the 405b model. Details: 🧵⬇️
1184
Sheridan Feucht @sfeucht.bsky.social · 30/11/2024
Love the Gabriel Garcia Marquez quote at the beginning. On my reading list!
010
Sheridan Feucht @sfeucht.bsky.social · 24/11/2024
Japonaise Bakery in Brookline :) 🥐
A box with anpan, melonpan, a strawberry croissant, and a matcha adzuki cream puff.
110
Reposted by Sheridan Feucht
Alex Makelov @amakelov.bsky.social · 24/11/2024
yes, this is what mechanistic interpretability research looks like
Cat sitting on a chair in front of a parked black car with its rear wheel removed and a hydraulic jack supporting it
2232