David Atkinson @diatkinson.bsky.social · 06/10/2026Read the paper: iii.baulab.info Joint work with Dillon Plunkett and @davidbau.bsky.social. 010
David Atkinson @diatkinson.bsky.social · 06/10/2026Many caveats! Some of them: this is a simple task using linear preferences over just 5 attributes; the test separates groups, not individual models; and we use LoRA adapters, rather than full fine-tunes. 120
David Atkinson @diatkinson.bsky.social · 06/10/2026We like this test because it doesn't rely on understanding the report. The model could answer in a language we don't speak, for example, and it would still work It complements concept-injection experiments like @Jack_W_Lindsey's, which test grounding by injecting known thoughts x.com/AnthropicAI...x.comAnthropic (@AnthropicAI) on XNew Anthropic research: Signs of introspection in LLMs. Can language models recognize their own internal thoughts? Or do they just make up plausible answers when asked about them? We found eviden… 130
David Atkinson @diatkinson.bsky.social · 06/10/2026We find that the cosine similarity between a model's decision and report attributions is 0.34 for faithful models compared to 0.08 for unfaithful ones (95% CI for the difference: 0.16 to 0.36). 110
David Atkinson @diatkinson.bsky.social · 06/10/2026To find out, we trained 32 new characters into each checkpoint, then used attribution patching to score every new adapter weight on each task. Each adapter in a pair was trained identically, differing only in the underlying base checkpoint. 110
David Atkinson @diatkinson.bsky.social · 06/10/2026Lots of work shows LLMs can describe behaviors they were only trained to perform (e.g. Owain Evans and co), and Josh Engels et al. (x.com/JoshAEngels...) traced one such case to a simple learned steering vector. Our question: can shared mechanisms tell faithful self-reports from unfaithful ones?x.comJosh Engels (@JoshAEngels) on X1/6: A recent paper shows that that LLMs are "self aware": when trained to exhibit a behavior like "risk taking", LLMs self report being risky. In a recent blog post, we explore what's happening h… 120
David Atkinson @diatkinson.bsky.social · 06/10/2026We can test this further: trained on all 40 layers, Qwen3-14B is a terrible self-reporter. But if we train only its first 20 layers, faithfulness reaches 0.74. Once training reaches layer 25 or beyond, faithfulness plummets, although the decisions are ~just as good. 120
David Atkinson @diatkinson.bsky.social · 06/10/2026What changed? Ablating adapter layers from the front or back shows that the faithful checkpoint stores its preferences 5-6 layers earlier. Our hypothesis: self-report only works once preferences sit early enough for the model's existing verbalization machinery to read them. 120
David Atkinson @diatkinson.bsky.social · 06/10/2026Although we train solely on decisions, faithful self-report emerges late in training, long after decisions have become accurate! Qwen3-32B at step 1000: decisions 0.82, faithfulness 0.25. At step 3000: decisions 0.92, faithfulness 0.83. This gives us a contrast pair. 120
David Atkinson @diatkinson.bsky.social · 06/10/2026Then, in a fresh context, we ask the model how it would weigh each attribute. This gives us two metrics: decision performance (how well the choices follow the character's hidden preferences) and faithfulness (how well the stated preferences match those revealed by its choices). 100
David Atkinson @diatkinson.bsky.social · 06/10/2026We build on Plunkett et al.'s "Self-Interpretability" setup (arxiv.org/abs/2505.17120): train Qwen3-32B to make decisions as 100 different characters (Gregor Samsa buying a washing machine...), each with random hidden preferences. 110
David Atkinson @diatkinson.bsky.social · 06/10/2026New COLM paper: Identifying Introspection From the Inside When an LLM tells us about its decisions, does it 𝘬𝘯𝘰𝘸 what drives its choices—or is it guessing? In our setting, we find that faithful models decide and report with the same layers. Unfaithful ones don't. 🧵 1103
Reposted by David AtkinsonNatalie Shapira @natalieshapira.bsky.social · 23/02/2026In this amazing multidisciplinary collaboration, we report our early experience with the @openclaw-x.bsky.social -> 14022
Reposted by David AtkinsonEric Todd @ericwtodd.bsky.social · 22/01/2026Can you solve this algebra puzzle? 🧩 cb=c, ac=b, ab=? A small transformer can learn to solve problems like this! And since the letters don't have inherent meaning, this lets us study how context alone imparts meaning. Here's what we found:🧵⬇️ 24811
Reposted by David AtkinsonArnab Sen Sharma @arnabsensharma.bsky.social · 04/11/2025How can a language model find the veggies in a menu? New pre-print where we investigate the internal mechanisms of LLMs when filtering on a list of options. Spoiler: turns out LLMs use strategies surprisingly similar to functional programming (think "filter" from python)! 🧵 1249
Reposted by David Atkinsonnikhil07prakash.bsky.social @nikhil07prakash.bsky.social · 24/06/2025How do language models track mental states of each character in a story, often referred to as Theory of Mind? We reverse-engineered how LLaMA-3-70B-Instruct handles a belief-tracking task and found something surprising: it uses mechanisms strikingly similar to pointer variables in C programming! 25918
Reposted by David AtkinsonDavid Bau @davidbau.bsky.social · 16/03/2025What will be the linchpin for AI dominance? Read our NSF/OSTP recommendations written with Goodfire's Tom McGrath tommcgrath.github.io, Transluce's Sarah Schwettmann cogconfluence.com, MIT's Dylan Hadfield-Menell @dhadfieldmenell.bsky.social TLDR; Dominance comes from **interpretability** 🧵 ↘️ 1218
Reposted by David AtkinsonChantal @chantalsh.bsky.social · 10/03/2025I'm searching for some comp/ling experts to provide a precise definition of “slop” as it refers to text (see: corp.oup.com/word-of-the-...) I put together a google form that should take no longer than 10 minutes to complete: forms.gle/oWxsCScW3dJU... If you can help, I'd appreciate your input! 🙏corp.oup.comOxford Word of the Year 2024 - Oxford University PressThe Oxford Word of the Year 2024 is 'brain rot'. Discover more about the winner, our shortlist, and 20 years of words that reflect the world. 0108
Reposted by David AtkinsonAndrew Lee @ajyl.bsky.social · 20/02/2025Excited about recent reasoning models? What is happening under the hood? Join ARBOR: Analysis of Reasoning Behaviors thru *Open Research* - a radically open collaboration to reverse-engineer reasoning models! Learn more: arborproject.github.io 1/Narborproject.github.ioARBOR 1133
Reposted by David AtkinsonEkdeep Singh @ ICML @ekdeepl.bsky.social · 16/02/2025New paper–accepted as *spotlight* at #ICLR2025! 🧵👇 We show a competition dynamic between several algorithms splits a toy model’s ICL abilities into four broad phases of train/test settings! This means ICL is akin to a mixture of different algorithms, not a monolithic ability. 2325
Reposted by David AtkinsonDavid Bau @davidbau.bsky.social · 07/12/2024PhD Applicants: remember that the Northeastern Computer Science PhD application deadline is Dec 15. It's a terrific time to do a PhD, with so many interesting things happening in AI. Apply here: www.khoury.northeastern.edu/apply/phd-ap...khoury.northeastern.eduPhD Apply - Khoury College of Computer Sciences 0335