Sign in

Chris Wendler

@wendlerc.bsky.social
559 followers 444 following 36 posts

Postdoc at the interpretable deep learning lab at Northeastern University, deep learning, LLMs, mechanistic interpretability

PostsRepliesMedia
Chris Wendler @wendlerc.bsky.social · 24/02/2026
I am not very disciplined about syncing my bluesky and x account, if you are interested what I am up to please check out my x account x.com/wendlerch or website wendlerc.github.io
wendlerc.github.io
Chris Wendler
030
Reposted by Chris Wendler
Clément Dumas @butanium.bsky.social · 30/06/2025
Our mech interp ICML workshop paper got accepted to ACL 2025 main! 🎉 In this updated version, we extended our results to several models and showed they can actually generate good definitions of mean concept representations across languages.🧵
x.com
Clément Dumas on X: "Excited to share our latest paper, accepted as a spotlight at the #ICML2024 mechanistic interpretability workshop! We find evidence that LLMs use language-agnostic representations of concepts 🧵↘️ https://t.co/dDS5iv199i" / X
Excited to share our latest paper, accepted as a spotlight at the #ICML2024 mechanistic interpretability workshop! We find evidence that LLMs use language-agnostic representations of concepts 🧵↘️ https://t.co/dDS5iv199i
191
Reposted by Chris Wendler
Can @canrager.bsky.social · 13/06/2025
Can we uncover the list of topics a language model is censored on? Refused topics vary strongly among models. Claude-3.5 vs DeepSeek-R1 refusal patterns:
1104
Reposted by Chris Wendler
Natalie Shapira @natalieshapira.bsky.social · 24/06/2025
I am really proud to share our work led by Nikhil Prakash and in collaboration with more mechanistic interpretability and Theory of Mind (ToM) researchers: arxiv.org/abs/2505.14685 You can find a tweet here with nice animations: x.com/nikhil07prak...
arxiv.org
Language Models use Lookbacks to Track Beliefs
How do language models (LMs) represent characters' beliefs, especially when those beliefs may differ from reality? This question lies at the heart of understanding the Theory of Mind (ToM) capabilitie...
092
Chris Wendler @wendlerc.bsky.social · 08/04/2025
Check out Sheridan’s work on concept induction circuits -- the soft version of induction we were promised a while ago :) During our multilingual concept patching experiments I have always been wondering whether it is those circuits doing the work. Finally, some evidence:
020
Chris Wendler @wendlerc.bsky.social · 21/03/2025
We also have a website sdxl-unbox.epfl.ch and a paper arxiv.org/abs/2410.22366
sdxl-unbox.epfl.ch
Unboxing SDXL Turbo with SAEs
Sparse Autoencoders (SAEs) find interpretable features in Stable Diffusion Turbo and enable fine-grained image editing.
000
Chris Wendler @wendlerc.bsky.social · 21/03/2025
Huge shoutout to Viacheslav Surkov who executed this project! This is what can happen when you keep pushing on your course project :P Really amazing!
100
Chris Wendler @wendlerc.bsky.social · 21/03/2025
If you go and play with this app you are guaranteed to find some fascinating quirk about SDXL turbo that no-one has ever seen before, which is why I love this work!
100
Chris Wendler @wendlerc.bsky.social · 21/03/2025
If you are someone who is great at designing user interfaces and want to build a better app or website with us, reach out to me via DM. If you are someone just curious about deep learning and diffusion models, go play with the features. We have more than 2000 features per layer.
110
Chris Wendler @wendlerc.bsky.social · 21/03/2025
It should be pretty self explanatory to use this app. You type in the feature index, select the layer, the strength of the coefficient, you brush a mask where the feature should be activated and hit "apply"...
100
Chris Wendler @wendlerc.bsky.social · 21/03/2025
You an also do more "abstract" things like brushing the face with a "water"-texture feature...
100
Chris Wendler @wendlerc.bsky.social · 21/03/2025
But that's not the best part yet. My favorite layer is the "style" layer. It allows you to draw with textures without modifying the rest of the image much. E.g. this happens when you brush the face with the "giraffe texture feature".
100
Chris Wendler @wendlerc.bsky.social · 21/03/2025
Inspired by this I also made one where I tried to take the hole-feature from a "Trypophobia" image...
100
Chris Wendler @wendlerc.bsky.social · 21/03/2025
Let's see what happens if we turn on a feature that activates on the beard but in the detail layer... We noticed in our experiments that these features often latch onto the context of the generated image (and require relevant context to be effective). The result is wild!
100
Chris Wendler @wendlerc.bsky.social · 21/03/2025
There is also one that seems to have something to do with the beard. Turning it on shows that it probably is more than just a beard... maybe a "manliness" feature or something like that.
100
Chris Wendler @wendlerc.bsky.social · 21/03/2025
We can look for interesting features in the "explore" tab. E.g. in the "composition" block feature number 199 seems to have to do with that hat. Let's turn it on...
100
Chris Wendler @wendlerc.bsky.social · 21/03/2025
Let's start with the prompt "an image of a colorful model"
100
Chris Wendler @wendlerc.bsky.social · 21/03/2025
We built an app where you can explore these features and turn them on while generating an image. The app is here: huggingface.co/spaces/surok...
huggingface.co
Unboxing SDXL with SAEs - a Hugging Face Space by surokpro2
Discover amazing ML apps made by the community
100
Chris Wendler @wendlerc.bsky.social · 21/03/2025
In case you ever wondered what you could do if you had SAEs for intermediate results of diffusion models, we trained SDXL Turbo SAEs on 4 blocks for you. We noticed that they specialize into a "composition", a "detail", and a "style" block. And one that is hard to make sense of.
120
Chris Wendler @wendlerc.bsky.social · 18/03/2025
Apply to Akhil's lab, he is great!
020
Reposted by Chris Wendler
Aaron Mueller @amuuueller.bsky.social · 11/03/2025
Lots of work coming soon to @iclr-conf.bsky.social and @naaclmeeting.bsky.social in April/May! Come chat with us about new methods for interpreting and editing LLMs, multilingual concept representations, sentence processing mechanisms, and arithmetic reasoning. 🧵
1206
Chris Wendler @wendlerc.bsky.social · 10/03/2025
And are more transferrable. I think with the current methods, we get the best interpretations by accumulating many different agreeing angles onto the same question. SAE's can be one of them, salience maps another one, patching another one, and so on. But we still need better tools.
000
Chris Wendler @wendlerc.bsky.social · 10/03/2025
I am also skeptical about SAEs and steering. I usually compare SAE/steering interventions to adversarial attacks (AAs). For AAs we know that they can yield arbitrary outputs via minimal perturbations of arbitrary intermediate states. Compared to those steering/SAE IMO use less optimization power.
sdxl-unbox.epfl.ch
Unboxing SDXL Turbo with SAEs
Sparse Autoencoders (SAEs) find interpretable features in Stable Diffusion Turbo and enable fine-grained image editing.
100
Chris Wendler @wendlerc.bsky.social · 10/03/2025
What exactly this tells us about the mechanisms is an open question. Compared to editing the input, you can get very different but still interpretable effects, e.g., in our work we basically found features that can be turned into style brushes by SAEing up.0.1 in SDXL Turbo sdxl-unbox.epfl.ch
sdxl-unbox.epfl.ch
Unboxing SDXL Turbo with SAEs
Sparse Autoencoders (SAEs) find interpretable features in Stable Diffusion Turbo and enable fine-grained image editing.
100
Chris Wendler @wendlerc.bsky.social · 10/03/2025
I did not read that paper with the random weights, but this is something I imagine impossible to do with random weights and thus probably not properly discussed in that work.
000
Chris Wendler @wendlerc.bsky.social · 10/03/2025
This is different angle that makes interpretations of SAE features testable. „Do they affect the remaining forward pass in a way consistent with my interpretation?“
200
Chris Wendler @wendlerc.bsky.social · 10/03/2025
So I agree that random NN features + learnt SAE is a powerful encoder and one has to be careful with the interpretation. But Clement was talking about interventions that in a generative model modify its output also accordingly.
100
Chris Wendler @wendlerc.bsky.social · 10/03/2025
I think that it makes total sense that learning SAE features, which is not much different from computing clusters (at least for k=1), on some layer of a randomly initialised NN should pick up interesting features of your data. People used random features as basis for ML since the beginning.
100
Chris Wendler @wendlerc.bsky.social · 10/03/2025
But I will say that SAE features often (especially with expansion factors common for the ones used in LLMs) are not good for interventions.
000
Chris Wendler @wendlerc.bsky.social · 10/03/2025
Honestly, I think Clément made a great point here and it is hard to get what you mean with this response.
210
Chris Wendler @wendlerc.bsky.social · 05/03/2025
Seems like you are not giving useful inputs tbh.
000
Chris Wendler @wendlerc.bsky.social · 03/03/2025
Yes: transluce.org/neuron-descr...
transluce.org
Scaling Automatic Neuron Description<!-- --> | Transluce AI
120
Reposted by Chris Wendler
Andrew Lee @ajyl.bsky.social · 20/02/2025
Excited about recent reasoning models? What is happening under the hood? Join ARBOR: Analysis of Reasoning Behaviors thru *Open Research* - a radically open collaboration to reverse-engineer reasoning models! Learn more: arborproject.github.io 1/N
arborproject.github.io
ARBOR
1133
Chris Wendler @wendlerc.bsky.social · 17/02/2025
This seems like an elegant idea!
020
Reposted by Chris Wendler
David Bau @davidbau.bsky.social · 31/01/2025
DeepSeek R1 shows how important it is to be studying the internals of reasoning models. Try our code: Here @canrager.bsky.social shows a method for auditing AI bias by probing the internal monologue. dsthoughts.baulab.info I'd be interested in your thoughts.
dsthoughts.baulab
1289
Reposted by Chris Wendler
Nathan Lambert @natolambert.bsky.social · 18/12/2024
The AI agent spectrum Separating different classes of AI agents from a long history of reinforcement learning. Why we can be optimistic for AI agents but also extremely critical of the terrible communications around them to date. Plus, some policy guidance.
buff.ly
The AI Agent Spectrum
Separating different classes of AI agents from a long history of reinforcement learning.
46610
Chris Wendler @wendlerc.bsky.social · 13/12/2024
The resources you find online on transformers are just next level... My jaw dropped when I first stumbled upon this video series: www.youtube.com/watch?v=V3NQ...
youtube.com
0L - Theory [rough early thoughts]
YouTube video by Mechanistic Interpretability
050
Chris Wendler @wendlerc.bsky.social · 10/12/2024
Not if you give them reliable information and a piece of paper.
010
Reposted by Chris Wendler
Alexander Kolesnikov @handle.invalid · 04/12/2024
Ok, it is yesterdays news already, but good night sleep is important. After 7 amazing years at Google Brain/DM, I am joining OpenAI. Together with @xzhai.bsky.social and @giffmana.ai, we will establish OpenAI Zurich office. Proud of our past work and looking forward to the future.
811611
Chris Wendler @wendlerc.bsky.social · 25/11/2024
bit grumpy but great summary of the tokenformer paper www.youtube.com/watch?v=gfU5...
youtube.com
TokenFormer: Rethinking Transformer Scaling with Tokenized Model Parameters (Paper Explained)
YouTube video by Yannic Kilcher
020
Chris Wendler @wendlerc.bsky.social · 23/11/2024
Does this mean that this thing is something like a bag of words (sum of word embeddings)?
110
Chris Wendler @wendlerc.bsky.social · 22/11/2024
This model2vec sounds interesting, I wonder how it works.
110
Reposted by Chris Wendler
Julian Minder @jkminder.bsky.social · 22/11/2024
Can we understand and control how language models balance context and prior knowledge? Our latest paper shows it’s all about a 1D knob! 🎛️ arxiv.org/abs/2411.07404 Co-led with @kevdududu.bsky.social - @niklasstoehr.bsky.social , Giovanni Monea, @wendlerc.bsky.social, Robert West & Ryan Cotterell.
1133
Chris Wendler @wendlerc.bsky.social · 20/11/2024
In case you also wondered how to derive the maximal update parametrisation (muP) learning rate for ADAM. I did a short write up: tinyurl.com/mup-for-adam. Thanks Ilia Badanin and Eugene Golikov for your help on this.
tinyurl.com
Notion – The all-in-one workspace for your notes, tasks, wikis, and databases.
A new tool that blends your everyday work apps into one. It's the all-in-one workspace for you and your team
062
Chris Wendler @wendlerc.bsky.social · 19/11/2024
Honestly I‘d recommend them to work through ARENA 3.0. I find that the most effective. www.arena.education/chapter1
arena.education
chapter1 — ARENA
250