Sign in

Sheridan Feucht

@sfeucht.bsky.social
252 followers 334 following 40 posts

PhD student doing LLM interpretability with @davidbau.bsky.social and @byron.bsky.social. (they/them) sfeucht.github.io

PostsRepliesMedia
Sheridan Feucht @sfeucht.bsky.social · 28/09/2026
[New preprint] How do images in VLMs align with words? In this work, we found a set of attention heads responsible for OCR. But to our surprise, these heads were actually able to verbalize much more than just text. So, we used them to create a simple logit lens for image tokens!
1107
Sheridan Feucht @sfeucht.bsky.social · 10/08/2026
Logging back on here just to make an ImageNet appreciation post. It's so refreshing to look at *real* random images instead of diffusion-generated ones. If you prompted FLUX to generate a photo of a desktop computer, you would never see that wallpaper... real images have some special quality to them
110
Reposted by Sheridan Feucht
Koyena Pal @koyena.bsky.social · 22/01/2026
Can models understand each other's reasoning? 🤔 When Model A explains its Chain-of-Thought (CoT) , do Models B, C, and D interpret it the same way? Our new preprint with @davidbau.bsky.social and @csinva.bsky.social explores CoT generalizability 🧵👇 (1/7)
1278
Sheridan Feucht @sfeucht.bsky.social · 22/01/2026
Eric’s setting is a really cool way of studying “pure” in-context learning, where symbols always mean different things in different contexts (so the model has to constantly do ICL). He has some really cool experiments you should check out 🔎
050
Sheridan Feucht @sfeucht.bsky.social · 14/12/2025
I just finished writing a blog post about a fake ICLR paper that caused a stir about a month ago. I think it has interesting implications for whether AI-generated text is "meaningful" or not. sfeucht.github.io/rerereading/
120
Sheridan Feucht @sfeucht.bsky.social · 06/12/2025
I'm headed to San Diego this afternoon to attend the mech interp workshop at #NeurIPS! Really excited to give a lightning talk and catch up with everyone. 🌴
041
Reposted by Sheridan Feucht
Kanishka Misra @kanishka.bsky.social · 28/07/2025
Looking forward to attending #cogsci2025 (Jul 29 - Aug 3)! I’m especially excited to meet students who will be applying to PhD programs in Computational Ling/CogSci in the coming cycle. Please reach out if you want to meet up and chat! Email is the best way, but DM also works if you must! quick🧵:
Placeholders for 3 students (number arbitrarily chosen) and me - to signify my eventual group!
1217
Sheridan Feucht @sfeucht.bsky.social · 22/07/2025
We've added a quick new section to this paper, which was just accepted to @COLM_conf! By summing weights of concept induction heads, we created a "concept lens" that lets you read out semantic information in a model's hidden states. 🔎
171
Reposted by Sheridan Feucht
Koyena Pal @koyena.bsky.social · 30/06/2025
🚨 Registration is live! 🚨 The New England Mechanistic Interpretability (NEMI) Workshop is happening Aug 22nd 2025 at Northeastern University! A chance for the mech interp community to nerd out on how models really work 🧠🤖 🌐 Info: nemiconf.github.io/summer25/ 📝 Register: forms.gle/v4kJCweE3UUH...
NEMI 2024 (Last Year)
0108
Sheridan Feucht @sfeucht.bsky.social · 24/06/2025
Nikhil's recent paper is a tour de force in causal analysis! They show that LLMs keep track of what characters know in a story using "pointer" mechanisms. Definitely worth checking out.
042
Sheridan Feucht @sfeucht.bsky.social · 25/04/2025
I used to think formal reasoning was central to language and intelligence, but now I’m not so sure. Wrote a short post about my thoughts on this, with a couple chewy anecdotes. Would love to get some feedback or pointers to further reading. sfeucht.github.io/syllogisms/
sfeucht.github.io
Sheridan Feucht
Solving Syllogisms is Not Intelligence April 23, 2025 (I think that we overvalue logical reasoning when it comes to measuring "intelligence.") What do we mean by intelligence in the context of cogniti...
160
Sheridan Feucht @sfeucht.bsky.social · 10/04/2025
I'll present a poster for this work at NENLP tomorrow! Come find me at poster #80...
071
Sheridan Feucht @sfeucht.bsky.social · 07/04/2025
[📄] Are LLMs mindless token-shifters, or do they build meaningful representations of language? We study how LLMs copy text in-context, and physically separate out two types of induction heads: token heads, which copy literal tokens, and concept heads, which copy word meanings.
17518
Reposted by Sheridan Feucht
Chantal @chantalsh.bsky.social · 10/03/2025
I'm searching for some comp/ling experts to provide a precise definition of “slop” as it refers to text (see: corp.oup.com/word-of-the-...) I put together a google form that should take no longer than 10 minutes to complete: forms.gle/oWxsCScW3dJU... If you can help, I'd appreciate your input! 🙏
corp.oup.com
Oxford Word of the Year 2024 - Oxford University Press
The Oxford Word of the Year 2024 is 'brain rot'. Discover more about the winner, our shortlist, and 20 years of words that reflect the world.
0108
Sheridan Feucht @sfeucht.bsky.social · 22/02/2025
I like this work a lot. Racism+misogyny in medicine is genuinely dangerous, so it's really important to keep tabs on model biases if we're going to use LLMs in clinical settings. It's nice to see that interpretability techniques are useful here.
060
Reposted by Sheridan Feucht
NDIF Team @ndif-team.bsky.social · 09/12/2024
Do you have a great experiment that you want to run on Llama 405b but not enough GPUs? 🚨 #NDIF is opening up more spots in our 405b pilot program! Apply now for a chance to conduct your own groundbreaking experiments on the 405b model. Details: 🧵⬇️
1184
Reposted by Sheridan Feucht
Alex Makelov @amakelov.bsky.social · 24/11/2024
yes, this is what mechanistic interpretability research looks like
Cat sitting on a chair in front of a parked black car with its rear wheel removed and a hydraulic jack supporting it
2232