Sign in

Benno Krojer

@bennokrojer.bsky.social
2.7K followers 1K following 1.8K posts

AI PhDing at Mila/McGill. Happily residing in Montreal 🥯❄️ Academic stuff: language grounding, vision+language, interp, rigorous & creative evals, cogsci Other: many sports, urban explorations, puzzles/quizzes bennokrojer.com

PostsRepliesMedia
Benno Krojer @bennokrojer.bsky.social · 05/06/2026
If I had to summarize this paper in one philosophical question: Why do we call some things by the same name and others by different names? Related to old philosophy of logic/language is Quine's work (1960): en.wikipedia.org/wiki/Inscrut...
en.wikipedia.org
Inscrutability of reference - Wikipedia
040
Benno Krojer @bennokrojer.bsky.social · 05/06/2026
Check out more examples in our demo: adadtur.github.io/nvrd-demo/ And here is the full paper: arxiv.org/abs/2606.05409
adadtur.github.io
NVRD Dataset Explorer
020
Benno Krojer @bennokrojer.bsky.social · 05/06/2026
My first last-author paper is out! If you saw this dog below and someone showed you the second image, would you consider them the same word/concept? (more examples in Ada's thread) We study if VLMs agree with humans on this and revisit old questions around shape vs. texture bias in vision
1124
Reposted by Benno Krojer
Elinor @elinorpd.bsky.social · 21/04/2026
I'll be presenting OvertonBench at #ICLR2026 in Rio later this week! 📍Sat, Apr 25, 10:30am in Pavilion 4 (#4109) Please DM me if you'd like to chat about pluralistic / value alignment, societal impacts, epistemology, fairness, evals, etc
081
Benno Krojer @bennokrojer.bsky.social · 21/04/2026
I'll be at ICLR! Will present our LatentLens paper at the Re-Align workshop and happy to chat (ideally at the beach 🏖️)
0100
Benno Krojer @bennokrojer.bsky.social · 31/03/2026
What are your favorite papers that can serve as excellent examples how to write great scientific paper, present results, great figures, make it engaging and easy to follow? Doesn't necessarily have to be the most cited or impactful ones
260
Reposted by Benno Krojer
Elinor @elinorpd.bsky.social · 10/03/2026
Inspired by @bennokrojer.bsky.social, we included a Behind the Scenes section 🎬 The goal is to make science more transparent 🔍, share lessons learned 🧠, and provide a more realistic lens on the research journey 👣 8/ bsky.app/profile/benn...
161
Reposted by Benno Krojer
Gaurav Kamath @grvkamath.bsky.social · 04/03/2026
🚨New Paper!🚨 How do reasoning LLMs handle inferences that have no deterministic answer? We find that they diverge from humans in some significant ways, and fail to reflect human uncertainty… 🧵(1/10)
35820
Benno Krojer @bennokrojer.bsky.social · 27/02/2026
*study in isolation
000
Benno Krojer @bennokrojer.bsky.social · 27/02/2026
True, but it's also more loopy! Not just one clean forward pass you can study
100
Benno Krojer @bennokrojer.bsky.social · 27/02/2026
Another way to put it: With ai systems, we'll never have such privileged access to latents than with brains (albeit just our own, so depends how much you think we're all the same roughly)
030
Benno Krojer @bennokrojer.bsky.social · 27/02/2026
The flip side: With ai systems it's still very much unclear which of our human intuitions apply (anthropomorphizing) and when they're a completely different beast that require fully new theories of cognition
000
Benno Krojer @bennokrojer.bsky.social · 27/02/2026
Sure, there's lots of fallacies and biases in that process of introspection but I wouldn't discard subjective experience as a very strong source of insight
020
Benno Krojer @bennokrojer.bsky.social · 27/02/2026
People often say (myself too): Interpretability on AI is so much easier than neuroscience! We can inspect everything and even retrain (vs carefully poke a little into the brain)! One big advantage in neuroscience I often forget: We're quite literally *inside* the thing we're studying
471
Benno Krojer @bennokrojer.bsky.social · 23/02/2026
This is a more accurate example since our method explicitly goes beyond single token interpretations --> words in the context of a sentence/paragraph
000
Benno Krojer @bennokrojer.bsky.social · 23/02/2026
For more detailed instructions how to use the library: github.com/McGill-NLP/l...
github.com
GitHub - McGill-NLP/latentlens: Code and data for the paper "LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs"
Code and data for the paper "LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs" - McGill-NLP/latentlens
000
Benno Krojer @bennokrojer.bsky.social · 23/02/2026
You can now "pip install latentlens" 🔨 It comes with: * pre-computed embeddings for several popular LLMs and VLMs * a txt file with sentences describing WordNet concepts, which we recommend as a standard corpus to get embeddings from * ... Try it out and let us know what we can improve!
272
Benno Krojer @bennokrojer.bsky.social · 23/02/2026
Is interpretability at the random fact-gathering stage or beyond?
020
Benno Krojer @bennokrojer.bsky.social · 23/02/2026
Finally getting into this classic Let's see if by the end I'll have a clearer idea what type of science some fields of AI are, like interpretability What are our paradigms?
270
Benno Krojer @bennokrojer.bsky.social · 20/02/2026
Google decided to show this as my first sentence from my website (and not any of the sentences actually at the top of the website)
010
Benno Krojer @bennokrojer.bsky.social · 16/02/2026
Keep me posted and feel free to ping me anytime something is confusing!
010
Benno Krojer @bennokrojer.bsky.social · 16/02/2026
Re 2) this was a typo and should be "i" for token position consistent with later uses in 3.2 and also how we use "i" in 3.1
000
Benno Krojer @bennokrojer.bsky.social · 16/02/2026
Maybe we can formulate it as a description d is text with optional meta-data (token position, layer) that is mapped to a vector r The general formalism is tricky but i think the intuition is hopefully clear :)
100
Benno Krojer @bennokrojer.bsky.social · 16/02/2026
So in our case (LatentLens) i would say: a description here is something like "a brown *dog*" and not "a brown dog" so the token position makes it a different description (this is also how we highlight it in our demo: bennokrojer.com/vlm_interp_d...)
bennokrojer.com
Image 0000 - LLaMA3-8B + ViT-L/14-336
100
Benno Krojer @bennokrojer.bsky.social · 16/02/2026
So I got a chance to look closely and you are right in both cases! Thank you for spotting this. I will upload a new version on arxiv soon with fixes To clarify things here also: 1) in 3.1 we described things generally but missed that eg LatentLens would match several vectors r with a description d
110
Benno Krojer @bennokrojer.bsky.social · 14/02/2026
Thank you! Let me get back to you later today on this when I'm on my laptop
100
Reposted by Benno Krojer
Vaibhav @vaibhavadlakha.bsky.social · 11/02/2026
What does it mean for visual tokens to be "interpretable" to LLM? And how to we measure it? These, and many more pressing questions are addressed! Introducing LatentLens -- a new, more faithful tool for interpretability! Honoured to have collaborated with @bennokrojer.bsky.social on this!
051
Benno Krojer @bennokrojer.bsky.social · 11/02/2026
Finally on a personal note, this will be the final paper of my PhD... what a journey it has been
010
Benno Krojer @bennokrojer.bsky.social · 11/02/2026
Pivoting to interpretability this year was great and i also wrote a blog post on this specifically: bennokrojer.com/interp.html
110
Benno Krojer @bennokrojer.bsky.social · 11/02/2026
This is a major lesson i will keep in mind for any future project: Test your assumptions, do not assume the field already has settled
110
Benno Krojer @bennokrojer.bsky.social · 11/02/2026
This project was definitely accelerated and shaped by Claude Code/Cursor. Building intuitive demos in interp is now much easier
110
Benno Krojer @bennokrojer.bsky.social · 11/02/2026
Finally we do test it empirically: finding some models where the embedding matrix of the LLM already provides decently interpretable nearest neighbors But this was not the full story yet... @mariusmosbach.bsky.social and @elinorpd.bsky.social nudged me to use contextual embeddings
111
Benno Krojer @bennokrojer.bsky.social · 11/02/2026
Then the project went "off-track" for a while, partially because we didn't question our assumptions enough: We just assumed visual tokens going into an LLM would not be that interpretable (based on the literature and our intuition) But we never fully tested it for many weeks!
110
Benno Krojer @bennokrojer.bsky.social · 11/02/2026
The initial ideation phase: Pivoting to a new direction, wondering what kind of interp work would be meaningful, getting feedback from my lab, ...
110
Benno Krojer @bennokrojer.bsky.social · 11/02/2026
For every one of my papers, I try to include a "Behind the Scenes" section I think this paper in particular has a lot going on behind the scenes; from lessons learned to personal reflections let me share some
130
Benno Krojer @bennokrojer.bsky.social · 11/02/2026
@delliott.bsky.social joined the project mid-way and somehow still had so much positive influence, ideas and energy. Good research is done with real care for detail and you can sense Des cares about the details
000
Benno Krojer @bennokrojer.bsky.social · 11/02/2026
I am very grateful to @sivareddyg.bsky.social's supervision in all these years, not just in challenging me to do impactful work, but also on the human side
100
Benno Krojer @bennokrojer.bsky.social · 11/02/2026
@mariusmosbach.bsky.social was an amazing mentor, his ideas and writing really shaped not just this work but how i conduct research
100
Benno Krojer @bennokrojer.bsky.social · 11/02/2026
This will be my last paper of the phd, can't believe it's been almost 5 years! It is the work i am most proud of and believe has the most potential. Feels right to wrap it up with this one
130
Benno Krojer @bennokrojer.bsky.social · 11/02/2026
We will keep working after this initial release, making latentlens as accessible as possible to other researchers and improving the code base We are optimistic latentlens can be used beyond visual inputs, and aim to make our codebase flexible for broader applications Share your feedback or ideas!
130
Benno Krojer @bennokrojer.bsky.social · 11/02/2026
Broader reflections: Are embedding spaces from diff modalities structurally similar, as the Platonic Representation Hypothesis suggests? Are LLMs so good at processing vision because pre-training induced an implicit physical world model? Multimodal interpretability is getting a bigger topic!
131
Benno Krojer @bennokrojer.bsky.social · 11/02/2026
Takeaways: We, the authors, were genuinely surprised to find such systematically high interpretability Recently people started using logit lens to study visual tokens in LLMs. We encourage the community to try out LatentLens next time, even beyond visual processing (any latent LLM representation)
140
Benno Krojer @bennokrojer.bsky.social · 11/02/2026
There are more cool analyses in the paper and the appendix that we encourage you to explore Or simply explore LatentLens and other tools in our interactive demo: bennokrojer.com/vlm_interp_... Teaser on some ablations we try: replace MLP with linear mapping, unfreeze LLM, worse training data, ...
130
Benno Krojer @bennokrojer.bsky.social · 11/02/2026
One last puzzle 🧩 How can LatentLens outperform EmbeddingLens even at layer 0? Our hypothsis: Visual tokens arrive already packaged in a semantic format Concretely: An input visual token might have the highest similarity with text representations at e.g. LLM layer 8 We call this "Mid-Layer Leap"
130
Benno Krojer @bennokrojer.bsky.social · 11/02/2026
Beyond our controlled setup, we also show how LatentLens works much better than baselines on off-the-shelf Qwen2-VL-7B-Instruct
131
Benno Krojer @bennokrojer.bsky.social · 11/02/2026
With this automatic metric, we can compare LogitLens, EmbeddingLens and LatentLens on 9 model combinations that we train (3 vision encoders x 3 LLMs) The two baselines are a mixed bag: some models and some layers are okay but many others not LatentLens shows high interpretability across the board
130
Benno Krojer @bennokrojer.bsky.social · 11/02/2026
How do we quantity whether a visual token is interpretable? We capture in a VLM judge what a human would intuitively do: Look at the top-5 NNs and the part in the image from which the visual token came from and answer: Are these top-5 NNs semantically related to the image or the part of the image?
130
Benno Krojer @bennokrojer.bsky.social · 11/02/2026
This might sound abstract, so here are some examples We can see: a) LatentLens makes visual tokens interpretable at all layers b) it has no issues with single-token results like logit lens and provides detailed full sentences! (no training or tuning involved)
130
Benno Krojer @bennokrojer.bsky.social · 11/02/2026
LatentLens in a nutshell: Instead of an LLM's fixed vocabulary (eg embedding matrix), we propose contextual representations as our pool of nearest neighbors → e.g. the representation of “dog” in “my brown dog” at some layer N
230
Benno Krojer @bennokrojer.bsky.social · 11/02/2026
Building a VLM can be surprisingly simple: You keep both the LLM and vision encoder frozen, you just train a small MLP that projects into the LLM embedding space as prefixes. That’s it 😮 But how and why does that work? How do visual tokens relate to language, i.e. do they have interpretable NNs?
141