Sign in

David Bau

@davidbau.bsky.social
2.4K followers 244 following 237 posts

Interpretable Deep Networks. baulab.info @davidbau

PostsRepliesMedia
David Bau @davidbau.bsky.social · 06/10/2026
The contrastive introspective setup is a great experimental platform. It illuminates how large-scale LMs might work so well. It also suggests a path for lie detection. Lots more cool things in this fascinating work. Well worth reading. iii.baulab.info/ bsky.app/profile/dia...
bsky.app
David Atkinson (@diatkinson.bsky.social)
Read the paper: iii.baulab.info Joint work with Dillon Plunkett and @davidbau.bsky.social.
040
David Bau @davidbau.bsky.social · 06/10/2026
It also teaches a key lesson. The same AI that can accurately say "I think X" on one hand can also be a stochastic parrot that lies about "I think X" other times. @diatkinson.bsky.social points out: we might be able to read the neurons to tell the difference. 🧵→ bsky.app/profile/dia...
bsky.app
David Atkinson (@diatkinson.bsky.social)
We find that the cosine similarity between a model's decision and report attributions is 0.34 for faithful models compared to 0.08 for unfaithful ones (95% CI for the difference: 0.16 to 0.36).
120
David Bau @davidbau.bsky.social · 06/10/2026
For the first time, @diatkinson.bsky.social is able to witness the neural footprint of faithful introspection. The introspective model stores knowledge in different neurons. It is direct evidence of a profound connection between introspection and generalization in LMs.🧵→ bsky.app/profile/dia...
bsky.app
David Atkinson (@diatkinson.bsky.social)
What changed? Ablating adapter layers from the front or back shows that the faithful checkpoint stores its preferences 5-6 layers earlier. Our hypothesis: self-report only works once preferences sit early enough for the model's existing verbalization machinery to read them.
140
David Bau @davidbau.bsky.social · 06/10/2026
What is AI doing when it introspects? The beauty is, since Qwen is an open model, we can now look inside to see what changes in the neural configuration. Compare the non-reflective earlier self and the introspective later self. What do we see? 🧵→ bsky.app/profile/dia...
bsky.app
David Atkinson (@diatkinson.bsky.social)
New COLM paper: Identifying Introspection From the Inside When an LLM tells us about its decisions, does it 𝘬𝘯𝘰𝘸 what drives its choices—or is it guessing? In our setting, we find that faithful models decide and report with the same layers. Unfaithful ones don't. 🧵
220
David Bau @davidbau.bsky.social · 06/10/2026
Then @diatkinson.bsky.social finds a breakthrough... It turns out, if you train Qwen LONGER on the A/B tasks, it can suddenly introspect. After just drilling it 3x longer, it somehow learns to write accurately about its own knowledge. This is nuts!! (?!?) 🧵→ bsky.app/profile/dia...
bsky.app
David Atkinson (@diatkinson.bsky.social)
Although we train solely on decisions, faithful self-report emerges late in training, long after decisions have become accurate! Qwen3-32B at step 1000: decisions 0.82, faithfulness 0.25. At step 3000: decisions 0.92, faithfulness 0.83. This gives us a contrast pair.
150
David Bau @davidbau.bsky.social · 06/10/2026
On one hand, failure to introspect is unsurprising. You train it to say A/B; why would you expect it to expound on thoughts? On the other hand many others have long noticed that frontier AI *CAN* introspect. Maybe Qwen is just too dumb to do it... 🧵 → bsky.app/profile/kma...
120
David Bau @davidbau.bsky.social · 06/10/2026
Next: ask it to introspect. What does Gregor like in washing machines? It will happily say "I think X". But it is TERRIBLE at it! Qwen acts as a stochastic parrot, spewing words about its thoughts that are NONSENSE. Unrelated to what it actually learned in A/B training. 🧵→
130
David Bau @davidbau.bsky.social · 06/10/2026
After just a bit of training Qwen totally gets it. Soon enough it can correctly answer A or B on new "Gregor Samsa washing machine puzzles". It generalizes. Unsurprising. Machine learning works. but ...🧵 →
120
David Bau @davidbau.bsky.social · 06/10/2026
First: teach Qwen something new. David uses SFT to train the LM to answer "A" or "B" on made-up trivia about famous people until the LM learns the trivia. Here it learns Gregor Samsa likes cheap washing machines even if they are noisy. Purely learning to say A or B. 🧵→ bsky.app/profile/dia...
bsky.app
David Atkinson (@diatkinson.bsky.social)
We build on Plunkett et al.'s "Self-Interpretability" setup (arxiv.org/abs/2505.17120): train Qwen3-32B to make decisions as 100 different characters (Gregor Samsa buying a washing machine...), each with random hidden preferences.
130
David Bau @davidbau.bsky.social · 06/10/2026
The craziest experiment in my lab is a new setup by @diatkinson.bsky.social where he has found a way induce introspection on open LMs using Dillon Plunkett's protocol, that lets him look inside introspection. It's nuts. Here's how the experiment works... 🧵 →
1101
Reposted by David Bau
Sheridan Feucht @sfeucht.bsky.social · 28/09/2026
[New preprint] How do images in VLMs align with words? In this work, we found a set of attention heads responsible for OCR. But to our surprise, these heads were actually able to verbalize much more than just text. So, we used them to create a simple logit lens for image tokens!
1107
David Bau @davidbau.bsky.social · 18/07/2026
You can play the contest too, at mazesofmenace.ai. Just point your coding agents at the repo github.com/davidbau/te... Nobody has cracked it yet: the field is wide open.
github.com
GitHub - davidbau/teleport-contest: The Teleport Coding Challenge — port NetHack 5.0 from C to JavaScript with bit-exact parity. Fork to enter.
The Teleport Coding Challenge — port NetHack 5.0 from C to JavaScript with bit-exact parity. Fork to enter. - davidbau/teleport-contest
220
David Bau @davidbau.bsky.social · 18/07/2026
Is it possible to write 100,000 lines of code well, if you do not read it? Let's go Hunting Zombies! davidbau.com/archives/20... In this post I dive into the code of two AI agent contestants in the Teleport coding challenge to learn their secrets. Very fun. And also instructive.
270
David Bau @davidbau.bsky.social · 17/06/2026
Mousing masks a few "gaze" attention heads to a spotlight and instantly shows the causal effects. Guide the model to the striped canopies, and then shift its attention to the autumn leaves, the fields of grain... Then check out the preprint! gaze.baulab.info
000
David Bau @davidbau.bsky.social · 17/06/2026
You can and read the preprint at gaze.baulab.info/#demo But first on that page scroll to the demo with WebGPU machine (most modern laptops+browsers) and click on "start demo". It's worth the wait as it downloads the 2B VLM.
100
David Bau @davidbau.bsky.social · 17/06/2026
Oh man! I love this preprint and also the website Rohit made to demo it. The gaze of a VLM is mediated by a much smaller set of attention heads than the full set, as if "conscious" attention is a small subset of "all attention heads". His demo lets you steer these in realtime.
180
David Bau @davidbau.bsky.social · 14/06/2026
Also check out the previous interview I had with Yascha about AI, which more of primer, here: writing.yaschamounk.com/p/david-bau
writing.yaschamounk.com
David Bau on How Artificial Intelligence Works
Yascha Mounk and David Bau delve into the “black box” of AI.
011
David Bau @davidbau.bsky.social · 14/06/2026
I recently spoke with Yascha Mounk about how researchers look inside AI to understand how it is thinking. Here is the podcast: writing.yaschamounk.com/p/david-bau-2
writing.yaschamounk.com
David Bau on How—and Whether—Artificial Intelligence Thinks
Yascha Mounk and David Bau examine the mysterious internal processes that drive AI behavior—and why they may be fundamentally alien.
162
David Bau @davidbau.bsky.social · 11/06/2026
Please join us at NEMI 2026, the 3rd New England Mechanistic Interpretability Workshop! August 14th at Boston University. Register now: nemiconf.github.io/summer26/ A remarkable time for AI. Come share your insights and research on the mechanisms inside our models. bsky.app/profile/mic...
bsky.app
Micah Benson (@micahben.bsky.social)
🧠🤖 The 2026 New England Mechanistic Interpretability (NEMI) Workshop will be Aug. 14 at Boston University! Help spread the word and join the New England mech interp community! Registration and submission info in thread:👇
163
David Bau @davidbau.bsky.social · 09/06/2026
Coding is not about performing a task! That's what ML is for. Code is about doing the *right* task. My prediction: as AI makes 777-levels of complexity routine, we will have much *more* code, not less. We will have an explosion of formal languages that, yes, humans will learn.
140
David Bau @davidbau.bsky.social · 09/06/2026
Musk has no ambition. A 777 has 2.6 million lines of code *not* because that's what it takes to fly. It's because that's what it takes to lift 370 tons safely from LAX to Heathrow every day without endangering its 400 passengers. www.youtube.com/watch?v=SZk...
youtube.com
Coding careers will die by Dec 2026: Elon Musk
Traditional coding as a job is standing on a burning platform. In...
150
David Bau @davidbau.bsky.social · 05/06/2026
📅 Key dates (Summer 2026): Applications close June 21 Competition begins June 29 Submissions due July 26 Awards event to invitees, Boston, Aug 25 Register: aletheias-quest.github.io Discord: discord.gg/4NX3sUC4ge
020
David Bau @davidbau.bsky.social · 05/06/2026
Run by Cadenza Labs and NDIF, compute supported by Schmidt Sciences, Coefficient Giving, and AWS. Submit code using NNsight API calls for white-box access during inference. You also get model weights for offline analysis and training, and benchmarks and samples to get started.
100
David Bau @davidbau.bsky.social · 05/06/2026
Your blue-team mission: submit code for a lie detector that catches the liars in the act. Your detector can use black-box or white-box inference: prompt, inspect, probe, patch, steer the target model's internals to catch it lying. Scored on unseen tests. You must generalize.
100
David Bau @davidbau.bsky.social · 05/06/2026
The contest is "Aletheia's Quest," a blue-team vs. red-team hackathon. Red teams (underway) have devised diverse scenarios that induce an AI to lie in specific instances. You enter as a blue team. $50K in prizes from Schmidt Sciences, with both white-box and black-box tracks. x.com/ndif_team/s...
x.com
NDIF (@ndif_team) on X
Aletheia (Ἀλήθεια) is the Greek spirit of truth. She stands against Lethe, the river of oblivion, and against the lie: the utterance that buries what's known. Her question, and ours: when a model speaks, does it speak the truth or conceal information?
131
David Bau @davidbau.bsky.social · 05/06/2026
We really don't understand AI lying fully. What happens when an AI lies? How can you best detect it? Can white-box detectors beat black-box ones in an adversarial setting? It is an ongoing interpretability challenge. We'd like you to try your hand at it. Join our contest!
110
David Bau @davidbau.bsky.social · 05/06/2026
But AI lie detection is hard and remains a central research challenge. Recent research suggests that simple probes can pick up on neural "tells" that reveal when it is lying, even when the output looks clean. anthropic.com/research/pr... arxiv.org/abs/2502.03407
arxiv.org
Detecting Strategic Deception Using Linear Probes
AI models might use deceptive strategies as part of scheming or misaligned behaviour. Monitoring outputs alone is insufficient, since the AI might produce seemingly benign outputs while their...
111
David Bau @davidbau.bsky.social · 05/06/2026
"You're right to call me on that!" Can you catch an AI in the act of lying? Register below to enter our AI lie-detection contest. AI lies are a big problem. The frontier labs have all worked hard to fight AI deception. They all try to monitor their AIs for it.
173
David Bau @davidbau.bsky.social · 06/05/2026
Just fork a repo to play. Tell your LLM agent to get the tests to pass. It comes with a starting code, detailed instructions, and a set of tests to score against. github.com/davidbau/te...
github.com
GitHub - davidbau/teleport-contest: The Teleport Coding Challenge — port NetHack 5.0 from C to JavaScript with bit-exact parity. Fork to enter.
The Teleport Coding Challenge — port NetHack 5.0 from C to JavaScript with bit-exact parity. Fork to enter. - davidbau/teleport-contest
010
David Bau @davidbau.bsky.social · 06/05/2026
The leaderboard is open. Click "play" to play one of the submitted NetHack ports, and "Tests" to see what points are being missed. mazesofmenace.ai/
100
David Bau @davidbau.bsky.social · 06/05/2026
In the contest, you can try your hand at steering the evolution of large-scale LLM "religion". I have created a porting sandbox, a skeleton start, a leaderboard, and visualization tools. I still do not know if agents can do it. Can you port NetHack 5?
100
David Bau @davidbau.bsky.social · 06/05/2026
The agents invented an idea, "sparse boundary frames" and defended it with rigorous authority. No such thing: They had created their own religion. With relentless efficiency, the agents infected 200K lines of code with their wrong ideas. Eventually I had to throw it all away.
100
David Bau @davidbau.bsky.social · 06/05/2026
LLMs are great at porting. Porting Hack and Rogue was easy: eighty-five minutes for Rogue. But then I pointed the agents at NetHack and watched them get stuck for three weeks at 18 failing tests. They were doing a hundred commits a day. But they couldn't move the bottom line.
100
David Bau @davidbau.bsky.social · 06/05/2026
Porting NetHack is fun because the 5.0 release this week is such an event: first major version in 10 years!!! But it is also interesting because NetHack is hard. I have spent four months trying to port it myself, with a swarm of AI agents, and the experience has humbled me. x.com/davidbau/st...
100
David Bau @davidbau.bsky.social · 06/05/2026
The Teleport Contest is open. Port NetHack 5.0 from C to JavaScript, bit-exactly. Same screen, every keystroke. Any approach: LLM agents, hand-coded, transpiler, hybrid. Live leaderboard, two phases through December. mazesofmenace.ai/announcement
260
David Bau @davidbau.bsky.social · 03/05/2026
We will also plan a 2nd round to give bonuses for high scorers that can update to a future NetHack 5.1 with minimal, readable code changes. Our question: can agents create hundreds of thousands of lines of code that humans would want to own afterwards?
010
David Bau @davidbau.bsky.social · 03/05/2026
You score points by producing a readable, maintainable JS version whose input-output behavior matches official NetHack 5.0 turn-by-turn on a held-out set of gameplay sequences. We will define a starter template and some public test sessions.
110
David Bau @davidbau.bsky.social · 03/05/2026
I am very curious if frontier orchestration systems like steve-yegge.bsky.social's Gastown or @ryancarson.com's Antfarm can crack it quickly. If there is interest, Alex Boruch⁠-⁠Gruszecki and I will flesh out contest rules this week and run a leaderboard.
110
David Bau @davidbau.bsky.social · 03/05/2026
The question: Can a swarm of LLM coding assistants enable a single person to work with a program of this scale and complexity? Can Claude, Codex, and Gemini understand the mammoth codebase well enough to convert it all to JS? Sure? How quickly? Anybody want to race?
110
David Bau @davidbau.bsky.social · 03/05/2026
Under active development over four decades and tracing its origins to Rogue (1980) and Hack (1982), the 442,901 lines of C and Lua in NetHack 5.0 have become a deeply interconnected system with many layers of intricate behavior.
110
David Bau @davidbau.bsky.social · 03/05/2026
NetHack is one of the most complex and longest-lived open source programs ever written, and after 46 years, v5.0 shipped today. www.nethack.org/common/inde... And ... it is a VERY cool large codebase to work with in the LLM era.
1103
David Bau @davidbau.bsky.social · 20/04/2026
2026 is a whirlwind year for AI. Underlying it all is the greatest scientific mystery of our age. How does a neural network think? I talked w Oliver Whang in NYTimes Magazine, on how AI interpretability is a tangle of structure waiting to be unraveled: www.nytimes.com/2026/04/15/...
2132
David Bau @davidbau.bsky.social · 08/04/2026
Tech industry mottos have a mixed track record. But we should hold idealists to their ideals. And we should celebrate when they come through. The Mythos non-release is a remarkable moment of conviction. Thoughts: davidbau.com/archives/20... Bravo to Anthropic's "race the top".
1133
David Bau @davidbau.bsky.social · 25/03/2026
Can you catch an AI lying? Red teams set up scenarios where models lie. Eg, do they lie under contextual pressure, even when not told to, but because honesty is costly? Then blue teams will build deception detectors using whitebox internals with NDIF. cadenza-labs.github.io/red-team-rfp/
031
David Bau @davidbau.bsky.social · 25/03/2026
Calling attention to an exciting "deception detection" hackathon we're planning this summer! w @NDIF and @CadenzaLabs. Recruiting red teams now, blue teams later. Red teams, time is short: proposals due Mar 31. $10K stipend + compute, $15K finals prize. nnsight.net/blog/2026/0...
252
David Bau @davidbau.bsky.social · 23/03/2026
Github here: github.com/davidbau/men...
Mazes of Menace, Github Readme page
010
David Bau @davidbau.bsky.social · 23/03/2026
The gap between Hack and NetHack taught me something about the future of CS. Worth thinking about. Does Computer Science Still Exist? davidbau.com/archives/20...
151
David Bau @davidbau.bsky.social · 23/03/2026
NetHack, the grandchild of Hack, is here: mazesofmenace.net/ I didn't write the code. Claude and Codex did. Rogue took 85 minutes. Hack took eight hours. NetHack (420,000 lines of C) has been grinding for two months and STILL isn't done!
110
David Bau @davidbau.bsky.social · 23/03/2026
Here is a 1982-rendition of the Logo programming language: mazesofmenace.net/logo/ Why? Because Logo was the community that connected all of them.
120
David Bau @davidbau.bsky.social · 23/03/2026
Rogue (1980) was the original that "Roguelikes" are like, by UCSC students Michael Toy and Glenn Wichman. Play it here: mazesofmenace.net/rogue/ Each one comes with a history of the people who made it.
110