Sign in

David Bau

@davidbau.bsky.social
2.3K followers 243 following 227 posts

Interpretable Deep Networks. baulab.info @davidbau

PostsRepliesMedia
Reposted by David Bau
Sheridan Feucht @sfeucht.bsky.social · 28/09/2026
[New preprint] How do images in VLMs align with words? In this work, we found a set of attention heads responsible for OCR. But to our surprise, these heads were actually able to verbalize much more than just text. So, we used them to create a simple logit lens for image tokens!
1107
David Bau @davidbau.bsky.social · 18/07/2026
Is it possible to write 100,000 lines of code well, if you do not read it? Let's go Hunting Zombies! davidbau.com/archives/20... In this post I dive into the code of two AI agent contestants in the Teleport coding challenge to learn their secrets. Very fun. And also instructive.
270
David Bau @davidbau.bsky.social · 17/06/2026
Oh man! I love this preprint and also the website Rohit made to demo it. The gaze of a VLM is mediated by a much smaller set of attention heads than the full set, as if "conscious" attention is a small subset of "all attention heads". His demo lets you steer these in realtime.
180
David Bau @davidbau.bsky.social · 14/06/2026
I recently spoke with Yascha Mounk about how researchers look inside AI to understand how it is thinking. Here is the podcast: writing.yaschamounk.com/p/david-bau-2
writing.yaschamounk.com
David Bau on How—and Whether—Artificial Intelligence Thinks
Yascha Mounk and David Bau examine the mysterious internal processes that drive AI behavior—and why they may be fundamentally alien.
162
David Bau @davidbau.bsky.social · 11/06/2026
Please join us at NEMI 2026, the 3rd New England Mechanistic Interpretability Workshop! August 14th at Boston University. Register now: nemiconf.github.io/summer26/ A remarkable time for AI. Come share your insights and research on the mechanisms inside our models. bsky.app/profile/mic...
bsky.app
Micah Benson (@micahben.bsky.social)
🧠🤖 The 2026 New England Mechanistic Interpretability (NEMI) Workshop will be Aug. 14 at Boston University! Help spread the word and join the New England mech interp community! Registration and submission info in thread:👇
163
David Bau @davidbau.bsky.social · 09/06/2026
Musk has no ambition. A 777 has 2.6 million lines of code *not* because that's what it takes to fly. It's because that's what it takes to lift 370 tons safely from LAX to Heathrow every day without endangering its 400 passengers. www.youtube.com/watch?v=SZk...
youtube.com
Coding careers will die by Dec 2026: Elon Musk
Traditional coding as a job is standing on a burning platform. In...
150
David Bau @davidbau.bsky.social · 05/06/2026
"You're right to call me on that!" Can you catch an AI in the act of lying? Register below to enter our AI lie-detection contest. AI lies are a big problem. The frontier labs have all worked hard to fight AI deception. They all try to monitor their AIs for it.
173
David Bau @davidbau.bsky.social · 06/05/2026
The Teleport Contest is open. Port NetHack 5.0 from C to JavaScript, bit-exactly. Same screen, every keystroke. Any approach: LLM agents, hand-coded, transpiler, hybrid. Live leaderboard, two phases through December. mazesofmenace.ai/announcement
260
David Bau @davidbau.bsky.social · 03/05/2026
NetHack is one of the most complex and longest-lived open source programs ever written, and after 46 years, v5.0 shipped today. www.nethack.org/common/inde... And ... it is a VERY cool large codebase to work with in the LLM era.
1103
David Bau @davidbau.bsky.social · 20/04/2026
2026 is a whirlwind year for AI. Underlying it all is the greatest scientific mystery of our age. How does a neural network think? I talked w Oliver Whang in NYTimes Magazine, on how AI interpretability is a tangle of structure waiting to be unraveled: www.nytimes.com/2026/04/15/...
2132
David Bau @davidbau.bsky.social · 08/04/2026
Tech industry mottos have a mixed track record. But we should hold idealists to their ideals. And we should celebrate when they come through. The Mythos non-release is a remarkable moment of conviction. Thoughts: davidbau.com/archives/20... Bravo to Anthropic's "race the top".
1133
David Bau @davidbau.bsky.social · 25/03/2026
Calling attention to an exciting "deception detection" hackathon we're planning this summer! w @NDIF and @CadenzaLabs. Recruiting red teams now, blue teams later. Red teams, time is short: proposals due Mar 31. $10K stipend + compute, $15K finals prize. nnsight.net/blog/2026/0...
252
David Bau @davidbau.bsky.social · 23/03/2026
In 1982, high school students in Sudbury, Mass. wrote a dungeon game called Hack. They had Atari 800s and Logo and an obsession with a Unix game called Rogue that most of them had never seen. I grew up one town over with the same computers and the same obsession.
171
David Bau @davidbau.bsky.social · 28/02/2026
Sam Altman and Dario Amodei have both staked out positions on AI weapons. But you can see from what they've said: the gap between them is a question of professional ethics. bsky.app/profile/mas...
bsky.app
Mike Masnick (@masnick.com)
My goodness.
150
David Bau @davidbau.bsky.social · 27/02/2026
I will be adding some time in my AI research group today for researchers and engineers to discuss the mission and ethics of all our work. We are often too preoccupied by the details. Good work requires clear purpose. Today is a good day to reflect.
090
David Bau @davidbau.bsky.social · 27/02/2026
Those of us who work in AI in the US today should take a moment to think today. Do not get distracted by the circus. Instead, let us pause to think carefully about our freedoms, our rights, and our responsibilities as citizens and professionals. It is a deadly serious moment.
1608
David Bau @davidbau.bsky.social · 23/02/2026
Are we all Agents of Chaos in AI? (Hope not!) In recent weeks using OpenClaw has taught us a lot about this wooly new kind of autonomous software agent. Its valuable to see what @NatalieShapira, @wendlerch et al. have seen: agentsofchaos.baulab.info/
2167
David Bau @davidbau.bsky.social · 21/02/2026
How do you knock the induction heads out of an LM while preserving its ability to think? Is it even possible? @keremsahin22.bsky.social's work is worth reading if you haven't seen it yet. hapax.baulab.info
1276
David Bau @davidbau.bsky.social · 27/01/2026
The Art of Wanting. About the question I see as central in AI ethics, interpretability, and safety. Can an AI take responsibility? I do not think so, but *not* because it's not smart enough. davidbau.com/archives/20...
193
Reposted by David Bau
Juan Diego Rodriguez @juand-r.bsky.social · 26/01/2026
I think everyone (not just academics) should read this.
061
David Bau @davidbau.bsky.social · 26/01/2026
What should academics be doing right now? I have been writing up some thoughts on what the research says about effective action, and what universities specifically can do. davidbau.github.io/poetsandnurs... It's on GitHub. Suggestions and pull requests welcome. github.com/davidbau/poe...
Federal agents with weapons drawn, moments before murdering American citizens on the streets of Minneapolis at the dawn of 2026.
03715
David Bau @davidbau.bsky.social · 25/01/2026
From induction to FVs, every ICL mechanism we've pinned down is fuzzy copying. Is copying all there is? @ericwtodd.bsky.social trained on groups where tokens have no fixed meaning and found a basket of mechanisms beyond copying. Watch them emerge, a grokking cascade! ↓ bsky.app/profile/eri...
060
David Bau @davidbau.bsky.social · 12/01/2026
I can't read Chinese, but my family has old genealogy documents I've always wanted to understand. Claude and Gemini helped me build an interactive reader to explore the calligraphy character by character. I can finally read my great-grandfather's epitaph. Try it: davidbau.com/archives/202...
Screenshot of Chinese calligraphy reader web application
0254
David Bau @davidbau.bsky.social · 06/01/2026
My vibe-coded Mandelbrot viewer is 40x faster now! New GPU synchronization tricks go outside the design intent of WebGPU specs. But the real story: Claude tells me what happens in the AGI break room. What superhuman AGIs say when the boss is not around: davidbau.com/archives/202...
030
David Bau @davidbau.bsky.social · 18/12/2025
I have been teaching myself to vibe code. Watch Claude Code grow my 780 lines to 13,600 - mandelbrot.page/coverage/ca... Two fundamental rules for staying in control: davidbau.com/archives/20...
2215
David Bau @davidbau.bsky.social · 11/12/2025
At the #Neurips2025 mechanistic interpretability workshop I gave a brief talk about Venetian glassmaking, since I think we face a similar moment in AI research today. Here is a blog post summarizing the talk: davidbau.com/archives/202...
The Doge of Venice visits a Murano glassworks in the 17th century. I will talk about why glassmaking in this era has some similarities to AI research today.
2195
David Bau @davidbau.bsky.social · 06/11/2025
The secret life of an LM is defined by its internal data types. Inner layers transport abstractions that are more robust than words, like concepts, functions, or pointers. In new work yesterday, @arnabsensharma.bsky.social et al identify a data type for *predicates*. bsky.app/profile/arn...
bsky.app
Arnab Sen Sharma (@arnabsensharma.bsky.social)
How can a language model find the veggies in a menu? New pre-print where we investigate the internal mechanisms of LLMs when filtering on a list of options. Spoiler: turns out LLMs use strategies surprisingly similar to functional programming (think "filter" from python)! 🧵
1152
David Bau @davidbau.bsky.social · 11/10/2025
What does an LLM do when it translates from Italian "amore" to Spanish "amor" or French "amour"? That's easy! (you might think) Because surely it knows: amore, amor, amour are all based on the same Latin word. It can just drop the "e", or add a "u".
2374
David Bau @davidbau.bsky.social · 06/10/2025
Looking forward to #COLM2025 tomorrow. DM me if you'll also be there and want to meet to chat.
050
David Bau @davidbau.bsky.social · 03/10/2025
There are a lot of interesting details that surface when you use SAEs to understand and control diffusion image synthesis models. Learn more in @wendlerc.bsky.social's talk.
140
David Bau @davidbau.bsky.social · 03/10/2025
On the Good Fight podcast w substack.com/@yaschamounk I give a quick but careful primer on how modern AI works. I also chat about our responsibility as machine learning scientists, and what we need to fix to get AI right. Take a listen and reshare - www.persuasion.community/p/david-bau
persuasion.community
David Bau on How Artificial Intelligence Works
Yascha Mounk and David Bau delve into the “black box” of AI.
073
David Bau @davidbau.bsky.social · 01/10/2025
I love the 'opinionated' approach taken by Aaron + team in this survey. It captures the ongoing work around the central casual puzzles we face in mechanistic interpretability.
030
David Bau @davidbau.bsky.social · 27/09/2025
Who is going to be at #COLM2025? I want to draw your attention to a COLM paper by my student @sfeucht.bsky.social that has totally changed the way I think and teach about LLM representations. The work is worth knowing. And you can meet Sheridan at COLM, Oct 7! bsky.app/profile/sfe...
1398
David Bau @davidbau.bsky.social · 26/09/2025
Announcing a broad expansion of the National Deep Inference Fabric. This could be relevant to your research...
1113
David Bau @davidbau.bsky.social · 20/09/2025
The NDIF youtube talk series continues... Don't miss the fascinating talks on by Xu Pan and Josh Engels, on the NDIF youtube channel. www.youtube.com/channel/UCaQ...
youtube.com
NDIF Team
We're a research computing project cracking open the mysteries inside large-scale AI systems. The NSF National Deep Inference Fabric consists of a unique combination of hardware and software that pr...
041
David Bau @davidbau.bsky.social · 20/09/2025
In the wake of the Jimmy Kimmel firing: Do not underestimate the power of the truth. The truth is our superpower. davidbau.com/archives/202...
davidbau.com
davidbau.com The Truth is Our Superpower
050
Reposted by David Bau
David Bau @davidbau.bsky.social · 29/08/2025
Monday: Trump tries to fire Fed Governor Lisa Cook (first time in 111 years). Thursday: CDC chief dismissed, four top scientists resign. Discredit, dismiss, blame. History shows exactly where this three-step pattern leads.
151
David Bau @davidbau.bsky.social · 29/08/2025
Monday: Trump tries to fire Fed Governor Lisa Cook (first time in 111 years). Thursday: CDC chief dismissed, four top scientists resign. Discredit, dismiss, blame. History shows exactly where this three-step pattern leads.
151
David Bau @davidbau.bsky.social · 18/08/2025
This Friday NEMI 2025 is at Northeastern in Boston, 8 talks, 24 roundtables, 90 posters; 200+ attendees. Thanks to goodfire.ai/ for sponsoring! nemiconf.github.io/summer25/ If you can't make it in person, the livestream will be here: www.youtube.com/live/4BJBis...
youtube.com
New England Mechanistic Interpretability Workshop
About:The New England Mechanistic Interpretability (NEMI) workshop aims to bring together academic and industry researchers from the New England and surround...
1167
David Bau @davidbau.bsky.social · 18/08/2025
Announcing a deep net interpretability talk series! Every week you will find new talks on recent research in the science of neural networks. The first few are posted: jackmerullo.bsky.social, Roy Rinberg, and me. At the @ndif-team.bsky.social Youtube Channel: www.youtube.com/@NDIFTeam
youtube.com
NDIF Team
We're a research computing project cracking open the mysteries inside large-scale AI systems. The NSF National Deep Inference Fabric consists of a unique combination of hardware and software that provides a remotely-accessible computing resource for scientists and students to perform detailed and reproducible experiments on large pretrained AI models, such as open large language models. We aim to make AI interpretability research more accessible through this channel by publishing lectures and educational content covering real interpretability research.
0115
David Bau @davidbau.bsky.social · 01/07/2025
The New England Mechanistic Interpretability Workshop, NEMI 2025 is August 22 in Boston. Talks, posters, meals, discussion... Most of all, an excellent chance to chat about new ideas with other great researchers in the field! Help spread the word - register and repost - bsky.app/profile/koy...
bsky.app
Koyena Pal (@koyena.bsky.social)
🚨 Registration is live! 🚨 The New England Mechanistic Interpretability (NEMI) Workshop is happening Aug 22nd 2025 at Northeastern University! A chance for the mech interp community to nerd out on how models really work 🧠🤖 🌐 Info: nemiconf.github.io/summer25/ 📝 Register: https://forms.gle/v4kJCweE3UUHUE81A
0122
David Bau @davidbau.bsky.social · 25/06/2025
The new "Lookback" paper from @nikhil07prakash.bsky.social‬ contains a surprising insight... 70b/405b LLMs use double pointers, akin to C programmers' double (**) pointers. They show up when the LLM is "knowing what Sally knows Ann knows", i.e., Theory of Mind. bsky.app/profile/nik...
bsky.app
@nikhil07prakash.bsky.social
How do language models track mental states of each character in a story, often referred to as Theory of Mind? We reverse-engineered how LLaMA-3-70B-Instruct handles a belief-tracking task and found something surprising: it uses mechanisms strikingly similar to pointer variables in C programming!
1293
David Bau @davidbau.bsky.social · 03/06/2025
FRIENDS: American science is being decimated by Congress NOW. Your help is needed to fix this. The current DC plan PERMANENTLY slashes NSF, NIH, all science training. Money isn't redirected—it's gone. Please read+share what's happening thevisible.net/posts/004-s...
150
David Bau @davidbau.bsky.social · 29/05/2025
Because of propaganda Americans do not understand what Rubio is doing with visas. "I gave you a visa to come and study," they think. x.com/CitizenFree... NO, he has not!! Please help explain to X how Rubio has stopped *ALL* student visas, and how it is killing US science.
160
David Bau @davidbau.bsky.social · 28/05/2025
When setting up my AI lab I faced a choice between Toronto and Boston. I chose Boston, my home and the world's best incubator for research talent. Here you can take a short stroll to meet with top minds in hundreds of fields from AI to astronomy, batteries to biotech.
1120
David Bau @davidbau.bsky.social · 25/05/2025
Black Box, Blood Money Friday evening, an Italian tourist escaped a torturer in Manhattan who was after his crypto password. I asked Anthropic's Opus 4 to analyze and explain what the episode might teach us about AI. It critiqued my guidance, instead proposing a focus on VCs:
220
David Bau @davidbau.bsky.social · 23/05/2025
Please join me in celebrating the contributions of our international students, researchers, and visitors. Here is a reminder of what makes America unique, and why no other nation can touch USA's 420 Nobel prizes:
4121
David Bau @davidbau.bsky.social · 17/05/2025
How to build AI leadership in the U.S? It is not about the chips. It is about the people! I spoke about AI interpretability at ntird.gov/ last week. (NTIRD is the joint program between 23 federal agencies that coordinates government technology investments.)
251
David Bau @davidbau.bsky.social · 10/05/2025
My grandfather was a WW2 American Army veteran who became a proud cataloger at the Library of Congress. As a toddler I remember walking to his Library office from his A Street home, with a stop at the playground. The LOC has always been the jewel of America for me.
1171
David Bau @davidbau.bsky.social · 03/05/2025
Leon Bottou's post ICLR thoughts are worth a read. He reminds us that modern AI is not just a product: it is a scientific wonder that we still do not understand. leon.bottou.org/news/two_les...
leon.bottou.org
news:two_lessons_from_iclr_2025 [leon.bottou.org]
1203