Sign in

Can

@canrager.bsky.social
85 followers 90 following 43 posts
PostsRepliesMedia
Can @canrager.bsky.social · 22/05/2026
The "tiling" perspective explains a lot of the common problems with SAEs for explaining concepts inside neural networks www.goodfire.ai/research/can...
020
Can @canrager.bsky.social · 07/05/2026
Goodfire has a new thread on neural geometry! I'm sold on steering along manifolds after this project. Agenda post: www.goodfire.ai/research/the... Paper on SAEs tiling manifolds: arxiv.org/abs/2604.28119 Paper on manifold steering: arxiv.org/abs/2605.05115
goodfire.ai
The World Inside Neural Networks
How neural geometry will unlock understanding and control of AI
020
Can @canrager.bsky.social · 28/02/2026
I've canceled my ChatGPT subscription.
0140
Can @canrager.bsky.social · 13/11/2025
Humans and LLMs think fast and slow. Do SAEs recover slow concepts in LLMs? Not really. Our Temporal Feature Analyzer discovers contextual features in LLMs, that detect event boundaries, parse complex grammar, and represent ICL patterns.
1208
Reposted by Can
Aaron Mueller @amuuueller.bsky.social · 01/10/2025
What's the right unit of analysis for understanding LLM internals? We explore in our mech interp survey (a major update from our 2024 ms). We’ve added more recent work and more immediately actionable directions for future work. Now published in Computational Linguistics!
24115
Reposted by Can
Koyena Pal @koyena.bsky.social · 30/06/2025
🚨 Registration is live! 🚨 The New England Mechanistic Interpretability (NEMI) Workshop is happening Aug 22nd 2025 at Northeastern University! A chance for the mech interp community to nerd out on how models really work 🧠🤖 🌐 Info: nemiconf.github.io/summer25/ 📝 Register: forms.gle/v4kJCweE3UUH...
NEMI 2024 (Last Year)
0108
Can @canrager.bsky.social · 13/06/2025
Can we uncover the list of topics a language model is censored on? Refused topics vary strongly among models. Claude-3.5 vs DeepSeek-R1 refusal patterns:
1104
Can @canrager.bsky.social · 20/02/2025
Announcing ARBOR, an open research community for collectively understanding how reasoning models like openai-o3 and deepseek-r1 work. We invite all researchers and enthusiasts to this initiative by @wattenberg.bsky.social's and @davidbau.bsky.social's lab. arborproject.github.io
161
Can @canrager.bsky.social · 29/01/2025
Addressing key concerns about AI competition. darioamodei.com/on-deepseek-...
darioamodei.com
Dario Amodei — On DeepSeek and Export Controls
On DeepSeek and Export Controls
011
Can @canrager.bsky.social · 10/01/2025
The #38c3 Chaos Computer Conference was a blast! 🚀 Find the accompanying code for my intro workshop on activation steering in the thread.
330
Can @canrager.bsky.social · 11/12/2024
Sparse Autoencoders (SAEs) are popular, with 10+ new approaches proposed in the last year. How do we know if we are making progress? The field has relied on imperfect proxy metrics. We are releasing SAE Bench, a suite of 8 SAE evaluations! Project co-led with Adam Karvonen.
142
Reposted by Can
NDIF Team @ndif-team.bsky.social · 10/12/2024
More big news! Applications are open for the NDIF Summer Engineering Fellowship—an opportunity to work on cutting-edge AI research infrastructure this summer in Boston! 🚀
196
Can @canrager.bsky.social · 09/12/2024
Safe travels to #NeurIPS2025 in Vancouver BC! Join our poster sessions on *Measuring Progress in Dictionary Learning with Board Game Models* and *Evaluating Sparse Autoencoders on Concept Erasure Tasks*. Reach out brainstorm future interpretability benchmarks.
031