Sign in

Thomas Fel

@thomasfel.bsky.social
1.4K followers 363 following 36 posts

Explainability, Computer Vision, Neuro-AI.🪴 Kempner Fellow @Harvard. Prev. PhD @Brown, @Google, @GoPro. Crêpe lover. 📍 Boston | 🔗 thomasfel.me

PostsRepliesMedia
Reposted by Thomas Fel
Kempner Institute at Harvard University @kempnerinstitute.bsky.social · 04/12/2025
Are you at #NeurIPS2025? Check out the #KempnerInstitute’s Day 2 presentations! 💡 #AI #NeuroAI @cpehlevan.bsky.social @kanakarajanphd.bsky.social @thomasfel.bsky.social @andykeller.bsky.social @binxuwang.bsky.social @njw.fish @yilundu.bsky.social
051
Reposted by Thomas Fel
Kempner Institute at Harvard University @kempnerinstitute.bsky.social · 12/11/2025
🐇Into the Rabbit Hull — Part 1: A Deep Dive into DINOv2🧠 Our latest Deeper Learning blog post is an #interpretability deep dive into one of today’s leading vision foundation models: DINOv2. 📖Read now: bit.ly/4nNfq8D Stay tuned — Part 2 coming soon. #AI #VLMs #DINOv2
bit.ly
Into the Rabbit Hull – Part I - Kempner Institute
This blog post offers an interpretability deep dive, examining the most important concepts emerging in one of today’s central vision foundation models, DINOv2. This blogpost is the first of a […]
1112
Thomas Fel @thomasfel.bsky.social · 06/11/2025
The Bau lab is on fire ! 😍
030
Reposted by Thomas Fel
Jennifer Hu @jennhu.bsky.social · 04/11/2025
Interested in doing a PhD at the intersection of human and machine cognition? ✨ I'm recruiting students for Fall 2026! ✨ Topics of interest include pragmatics, metacognition, reasoning, & interpretability (in humans and AI). Check out JHU's mentoring program (due 11/15) for help with your SoP 👇
02815
Reposted by Thomas Fel
Ida Momennejad @neuroai.bsky.social · 27/10/2025
Pleased to share new work with @sflippl.bsky.social @eberleoliver.bsky.social @thomasmcgee.bsky.social & undergrad interns at Institute for Pure and Applied Mathematics, UCLA. Algorithmic Primitives and Compositional Geometry of Reasoning in Language Models www.arxiv.org/pdf/2510.15987 🧵1/n
17715
Reposted by Thomas Fel
Thomas Serre @thomasserre.bsky.social · 24/10/2025
🧠 Thrilled to share our NeuroView with Ellie Pavlick! "From Prediction to Understanding: Will AI Foundation Models Transform Brain Science?" AI foundation models are coming to neuroscience—if scaling laws hold, predictive power will be unprecedented. But is that enough? Thread 🧵👇
2228
Reposted by Thomas Fel
Naomi Saphra @nsaphra.bsky.social · 16/10/2025
This is so cool. When you look at representational geometry, it seems intuitive that models are combining convex regions of "concepts", but I wouldn't have expected that this is PROVABLY true for attention or that there was such a rich theory for this kind of geometry.
2335
Thomas Fel @thomasfel.bsky.social · 15/10/2025
🕳️🐇Into the Rabbit Hull – Part II Continuing our interpretation of DINOv2, the second part of our study concerns the *geometry of concepts* and the synthesis of our findings toward a new representational *phenomenology*: the Minkowski Representation Hypothesis
2339
Thomas Fel @thomasfel.bsky.social · 14/10/2025
🕳️🐇 𝙄𝙣𝙩𝙤 𝙩𝙝𝙚 𝙍𝙖𝙗𝙗𝙞𝙩 𝙃𝙪𝙡𝙡 – 𝙋𝙖𝙧𝙩 𝙄 (𝑃𝑎𝑟𝑡 𝐼𝐼 𝑡𝑜𝑚𝑜𝑟𝑟𝑜𝑤) 𝗔𝗻 𝗶𝗻𝘁𝗲𝗿𝗽𝗿𝗲𝘁𝗮𝗯𝗶𝗹𝗶𝘁𝘆 𝗱𝗲𝗲𝗽 𝗱𝗶𝘃𝗲 𝗶𝗻𝘁𝗼 𝗗𝗜𝗡𝗢𝘃𝟮, one of vision’s most important foundation models. And today is Part I, buckle up, we're exploring some of its most charming features. :)
23612
Reposted by Thomas Fel
Meenakshi Khosla @meenakshikhosla.bsky.social · 08/10/2025
Superposition has reshaped interpretability research. In our @unireps.bsky.social paper led by @andre-longon.bsky.social we show it also matters for measuring alignment! Two systems can represent the same features yet appear misaligned if those features are mixed differently across neurons.
arxiv.org
Superposition disentanglement of neural representations reveals hidden alignment
The superposition hypothesis states that a single neuron within a population may participate in the representation of multiple features in order for the population to represent more features than the ...
292
Reposted by Thomas Fel
Jessica Hullman @jessicahullman.bsky.social · 09/10/2025
For XAI it’s often thought explanations help (boundedly rational) user “unlock” info in features for some decision. But no one says this, they say vaguer things like “supporting trust”. We lay out some implicit assumptions that become clearer when you take a formal view here arxiv.org/abs/2506.22740
arxiv.org
Explanations are a means to an end
Modern methods for explainable machine learning are designed to describe how models map inputs to outputs--without deep consideration of how these explanations will be used in practice. This paper arg...
2303
Reposted by Thomas Fel
David Picard @davidpicard.eurosky.social · 08/10/2025
🚨Updated: "How far can we go with ImageNet for Text-to-Image generation?" TL;DR: train a text2image model from scratch on ImageNet only and beat SDXL. Paper, code, data available! Reproducible science FTW! 🧵👇 📜 arxiv.org/abs/2502.21318 💻 github.com/lucasdegeorg... 💽 huggingface.co/arijitghosh/...
14410
Reposted by Thomas Fel
Greta Tuckute @gretatuckute.bsky.social · 04/10/2025
Check out @mryskina.bsky.social's talk and poster at COLM on Tuesday—we present a method to identify 'semantically consistent' brain regions (responding to concepts across modalities) and show that more semantically consistent brain regions are better predicted by LLMs.
0154
Reposted by Thomas Fel
Deniz Bayazit @bayazitdeniz.bsky.social · 25/09/2025
1/🚨 New preprint How do #LLMs’ inner features change as they train? Using #crosscoders + a new causal metric, we map when features appear, strengthen, or fade across checkpoints—opening a new lens on training dynamics beyond loss curves & benchmarks. #interpretability
2156
Reposted by Thomas Fel
Leshem (Legend) Choshen @EMNLP @lchoshen.bsky.social · 26/09/2025
Employing mechanistic interpretability to study how models learn, not just where they end up 2 papers find: There are phase transitions where features emerge and stay throughout learning 🤖📈🧠 alphaxiv.org/pdf/2509.17196 @amuuueller.bsky.social @abosselut.bsky.social alphaxiv.org/abs/2509.05291
261
Reposted by Thomas Fel
Tiago Pimentel @tpimentel.bsky.social · 14/07/2025
Mechanistic interpretability often relies on *interventions* to study how DNNs work. Are these interventions enough to guarantee the features we find are not spurious? No!⚠️ In our new paper, we show many mech int methods implicitly rely on the linear representation hypothesis🧵
Paper title "The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?" with the paper's graphical abstract showing how more powerful alignment maps between a DNN and an algorithm allow more complex features to be found and more "accurate" abstractions.
16612
Reposted by Thomas Fel
Sam Gershman @gershbrain.bsky.social · 28/09/2025
I was part of an interesting panel discussion yesterday at an ARC event. Maybe everybody knows this already, but I was quite surprised by how "general" intelligence was conceptualized in relation to human intelligence and the ARC benchmarks.
2243
Thomas Fel @thomasfel.bsky.social · 28/09/2025
Phenomenology → principle → method. From observed phenomena in representations (conditional orthogonality) we derive a natural instantiation. And it turns out to be an old friend: Matching Pursuit! 📄 arxiv.org/abs/2506.03093 See you in San Diego, @neuripsconf.bsky.social 🎉 #interpretability
arxiv.org
From Flat to Hierarchical: Extracting Sparse Representations with Matching Pursuit
Motivated by the hypothesis that neural network representations encode abstract, interpretable features as linearly accessible, approximately orthogonal directions, sparse autoencoders (SAEs) have bec...
050
Reposted by Thomas Fel
Malcolm Campbell @malcolmgcampbell.bsky.social · 19/09/2025
🚨Our preprint is online!🚨 www.biorxiv.org/content/10.1... How do #dopamine neurons perform the key calculations in reinforcement #learning? Read on to find out more! 🧵
1119871
Reposted by Thomas Fel
Isabel Papadimitriou @isabelpapad.bsky.social · 17/09/2025
Are there conceptual directions in VLMs that transcend modality? Check out our COLM oral spotlight 🔦 paper! We use SAEs to analyze the multimodality of linear concepts in VLMs with @chloesu07.bsky.social, @thomasfel.bsky.social, @shamkakade.bsky.social and Stephanie Gil arxiv.org/abs/2504.11695
1256
Thomas Fel @thomasfel.bsky.social · 17/09/2025
Check out our COLM 2025 (oral) 🎤 SAEs reveal that VLM embedding spaces aren’t just "image vs. text" cones. They contain stable conceptual directions, some forming surprising bridges across modalities. arxiv.org/abs/2504.11695 Demo 👉 vlm-concept-visualization.com
150
Reposted by Thomas Fel
Jennifer Hu @jennhu.bsky.social · 16/07/2025
Excited to announce the first workshop on CogInterp: Interpreting Cognition in Deep Learning Models @ NeurIPS 2025! 📣 How can we interpret the algorithms and representations underlying complex behavior in deep learning models? 🌐 coginterp.github.io/neurips2025/ 1/4
coginterp.github.io
Home
First Workshop on Interpreting Cognition in Deep Learning Models (NeurIPS 2025)
15819
Reposted by Thomas Fel
Andrew Lampinen @lampinen.bsky.social · 02/05/2025
How do language models generalize from information they learn in-context vs. via finetuning? In arxiv.org/abs/2505.00661 we show that in-context learning can generalize more flexibly, illustrating key differences in the inductive biases of these modes of learning — and ways to improve finetuning. 1/
arxiv.org
47722
Reposted by Thomas Fel
Harry Thasarathan @hthasarathan.bsky.social · 01/05/2025
Our work finding universal concepts in vision models is accepted at #ICML2025!!! My first major conference paper with my wonderful collaborators and friends @matthewkowal.bsky.social @thomasfel.bsky.social @Julian_Forsyth @csprofkgd.bsky.social Working with y'all is the best 🥹 Preprint ⬇️!!
0164
Reposted by Thomas Fel
Kosta Derpanis @csprofkgd.bsky.social · 01/05/2025
Accepted at #ICML2025! Check out the preprint. HUGE shoutout to Harry (1st PhD paper, in 1st year), Julian (1st ever, done as an undergrad), Thomas and Matt! @hthasarathan.bsky.social @thomasfel.bsky.social @matthewkowal.bsky.social
2357
Reposted by Thomas Fel
Gilles Louppe @glouppe.bsky.social · 29/04/2025
<proud advisor> Hot off the arXiv! 🦬 "Appa: Bending Weather Dynamics with Latent Diffusion Models for Global Data Assimilation" 🌍 Appa is our novel 1.5B-parameter probabilistic weather model that unifies reanalysis, filtering, and forecasting in a single framework. A thread 🧵
25215
Reposted by Thomas Fel
Arno Solin @arnosolin.bsky.social · 29/04/2025
Have you thought that in computer memory model weights are given in terms of discrete values in any case. Thus, why not do probabilistic inference on the discrete (quantized) parameters. @trappmartin.bsky.social is presenting our work at #AABI2025 today. [1/3]
34411
Reposted by Thomas Fel
Kempner Institute at Harvard University @kempnerinstitute.bsky.social · 28/04/2025
New in the Deeper Learning blog: Kempner researchers show how VLMs speak the same semantic language across images and text. bit.ly/KempnerVLM by @isabelpapad.bsky.social ,Chloe Huangyuan Su, @thomasfel.bsky.social, Stephanie Gil, and @shamkakade.bsky.social #AI #ML #VLMs #SAEs
bit.ly
Interpreting the Linear Structure of Vision-Language Model Embedding Spaces - Kempner Institute
Using sparse autoencoders, the authors show that vision-language embeddings boil down to a small, stable dictionary of single-modality concepts that snap together into cross-modal bridges. This resear...
093
Reposted by Thomas Fel
Martin Vinck @martinavinck.bsky.social · 10/04/2025
Firing rates in visual cortex show representational drift, while temporal spike sequences remain stable www.sciencedirect.com/science/arti... Great work by Boris Sotomayor and with @battaglialab.bsky.social
sciencedirect.com
Firing rates in visual cortex show representational drift, while temporal spike sequences remain stable
Neural firing-rate responses to sensory stimuli show progressive changes both within and across sessions, raising the question of how the brain mainta…
07020
Reposted by Thomas Fel
Greta Tuckute @gretatuckute.bsky.social · 10/04/2025
PINEAPPLE, LIGHT, HAPPY, AVALANCHE, BURDEN Some of these words are consistently remembered better than others. Why is that? In our paper, just published in J. Exp. Psychol., we provide a simple Bayesian account and show that it explains >80% of variance in word memorability: tinyurl.com/yf3md5aj
tinyurl.com
APA PsycNet
13914
Reposted by Thomas Fel
Satpreet (Sat) Singh @satpreetsingh.bsky.social · 07/04/2025
📽️Recordings from our @cosynemeeting.bsky.social #COSYNE2025 workshop on “Agent-Based Models in Neuroscience: Complex Planning, Embodiment, and Beyond" are now online: neuro-agent-models.github.io 🧠🤖
neuro-agent-models.github.io
🤖 Agent-Based Models in Neuroscience
13711
Thomas Fel @thomasfel.bsky.social · 07/03/2025
[...] overall, we argue an SAE does not just reveal concepts—it determines what can be seen at all." We propose to examine how constraints on SAE impose dual assumptions on the data, led by the amazing @sumedh-hindupur.bsky.social 😎
061
Reposted by Thomas Fel
Ekdeep Singh @ ICML @ekdeepl.bsky.social · 16/02/2025
New paper–accepted as *spotlight* at #ICLR2025! 🧵👇 We show a competition dynamic between several algorithms splits a toy model’s ICL abilities into four broad phases of train/test settings! This means ICL is akin to a mixture of different algorithms, not a monolithic ability.
2325
Reposted by Thomas Fel
TimDarcet @timdarcet.bsky.social · 14/02/2025
Want strong SSL, but not the complexity of DINOv2? CAPI: Cluster and Predict Latents Patches for Improved Masked Image Modeling.
14910
Reposted by Thomas Fel
Badr AlKhamissi @bkhmsi.bsky.social · 19/12/2024
🚨 New Paper! Can neuroscience localizers uncover brain-like functional specializations in LLMs? 🧠🤖 Yes! We analyzed 18 LLMs and found units mirroring the brain's language, theory of mind, and multiple demand networks! w/ @gretatuckute.bsky.social, @abosselut.bsky.social, @mschrimpf.bsky.social 🧵👇
210325
Reposted by Thomas Fel
Dmytro Mishkin @ducha-aiki.bsky.social · 08/02/2025
Scaling Laws in Patchification: An Image Is Worth 50,176 Tokens And More Feng Wang, Yaodong Yu, Guoyizhe Wei, Wei Shao, Yuyin Zhou, Alan Yuille, Cihang Xie tl;dr: we trained 1px patch size ViT so you don't have to. It improves results, but costly. arxiv.org/abs/2502.03738
1191
Thomas Fel @thomasfel.bsky.social · 07/02/2025
I'm delighted to share this latest research, led by the talented @hthasarathan.bsky.social and Julian. Their work uncovered both universal conceptual across models but also unique concepts specific to DINOv2 and SigLip! 🔥
040
Reposted by Thomas Fel
Harry Thasarathan @hthasarathan.bsky.social · 07/02/2025
🌌🛰️🔭Wanna know which features are universal vs unique in your models and how to find them? Excited to share our preprint: "Universal Sparse Autoencoders: Interpretable Cross-Model Concept Alignment"! arxiv.org/abs/2502.03714 (1/9)
15617
Reposted by Thomas Fel
Kristina Ulicna @kristinaulicna.bsky.social · 15/12/2024
Using mechanistic #interpretability 💻 to advance scientific #discovery 🧪 & capture striking biology? 🧬 Come see @jhartford.bsky.social's oral presentation 👨‍🏫 @ #NeurIPS2024 Interpretable AI workshop 🦾 to learn more about extracting features from large 🔬 MAEs! Paper 📄 ➡️: openreview.net/forum?id=jYl...
1222
Reposted by Thomas Fel
Remi Cadene @remicadene.bsky.social · 03/02/2025
Did you know that @PyTorch implements the Bessel's correction to standard deviation, but not numpy or jax. A possible source of disagreements when porting models to pytorch! @numpy_team
1213
Reposted by Thomas Fel
Ken Miller @kenmiller.bsky.social · 31/01/2025
New preprint: "The geometry of the neural state space of decisions", work by Mauro Monsalve-Mercado, buff.ly/42wVHD5. Surprising results & predictions! (Thread) We analyze neuropixel population recordings in macaque area LIP during a reaction time, random-dot motion 1/
Picture of neural manifolds for the two choices in a decision-making task, depicted in 3D and in 2D
317367
Reposted by Thomas Fel
Valérie Castin @vcastin.bsky.social · 31/01/2025
How do tokens evolve as they are processed by a deep Transformer? With José A. Carrillo, @gabrielpeyre.bsky.social and @pierreablin.bsky.social, we tackle this in our new preprint: A Unified Perspective on the Dynamics of Deep Transformers arxiv.org/abs/2501.18322 ML and PDE lovers, check it out!
29616
Reposted by Thomas Fel
AI Firehose @ai-firehose.column.social · 29/01/2025
The research introduces AXBENCH, demonstrating that prompting outperforms complex representation methods like sparse autoencoders. It features a new weakly-supervised method, ReFT-r1, which combines interpretability with competitive performance. arxiv.org/abs/2501.17148
arxiv.org
AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
ArXiv link for AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
021
Reposted by Thomas Fel
Kempner Institute at Harvard University @kempnerinstitute.bsky.social · 30/01/2025
The application for our #KempnerInstitute graduate fellowship for Harvard Ph.D. students is now open! Register for an upcoming virtual open house (2/6 & 2/18) and apply by 3/1. Details: bit.ly/40RY6qS
bit.ly
Graduate Fellowship - Kempner Institute
Kempner Graduate Fellowships support a dynamic and diverse community of PhD students across a number of graduate programs at Harvard, seeding new and innovative scientific discoveries in labs across t...
062
Thomas Fel @thomasfel.bsky.social · 30/01/2025
DinoV2, C:5232... 😶‍🌫️
210
Reposted by Thomas Fel
Dorsa Amir @dorsaamir.bsky.social · 25/01/2025
Does the culture you grow up in shape the way you see the world? In a new Psych Review paper, @chazfirestone.bsky.social & I tackle this centuries-old question using the Müller-Lyer illusion as a case study. Come think through one of history's mysteries with us🧵(1/13):
331093423
Reposted by Thomas Fel
Dr Abeba Birhane @abeba.blacksky.app · 20/01/2025
Join us at our upcoming workshop at ICLR, XAI4Science: From Understanding Model Behavior to Discovering New Scientific Knwlg, Apr 27-28 submissions on a-priori (ante-hoc) & a-posteriori (post-hoc) interpretability & self-explainable models for understanding model’s behvr welcm tinyurl.com/3w8sddpm
openreview.net
ICLR 2025 Workshop XAI4Science
Welcome to the OpenReview homepage for ICLR 2025 Workshop XAI4Science
13315
Reposted by Thomas Fel
Gabriel Peyré @gabrielpeyre.bsky.social · 22/01/2025
The Mathematics of Artificial Intelligence: In this introductory and highly subjective survey, aimed at a general mathematical audience, I showcase some key theoretical concepts underlying recent advancements in machine learning. arxiv.org/abs/2501.10465
214943
Reposted by Thomas Fel
Shiry Ginosar @shiryginosar.bsky.social · 09/01/2025
New paper! A self-supervised object-centric 2.1D image representation using 3D Gaussians, extending MAE with a Gaussian bottleneck. While Gaussian splatting has been used for single-scene reconstruction, we’re the first to apply it to image representation learning! brjathu.github.io/gmae/.
2375