David Atkinson @diatkinson.bsky.social · 06/10/2026New COLM paper: Identifying Introspection From the Inside When an LLM tells us about its decisions, does it 𝘬𝘯𝘰𝘸 what drives its choices—or is it guessing? In our setting, we find that faithful models decide and report with the same layers. Unfaithful ones don't. 🧵 192
Reposted by David AtkinsonNatalie Shapira @natalieshapira.bsky.social · 23/02/2026In this amazing multidisciplinary collaboration, we report our early experience with the @openclaw-x.bsky.social -> 14022
Reposted by David AtkinsonEric Todd @ericwtodd.bsky.social · 22/01/2026Can you solve this algebra puzzle? 🧩 cb=c, ac=b, ab=? A small transformer can learn to solve problems like this! And since the letters don't have inherent meaning, this lets us study how context alone imparts meaning. Here's what we found:🧵⬇️ 24811
Reposted by David AtkinsonArnab Sen Sharma @arnabsensharma.bsky.social · 04/11/2025How can a language model find the veggies in a menu? New pre-print where we investigate the internal mechanisms of LLMs when filtering on a list of options. Spoiler: turns out LLMs use strategies surprisingly similar to functional programming (think "filter" from python)! 🧵 1249
Reposted by David Atkinsonnikhil07prakash.bsky.social @nikhil07prakash.bsky.social · 24/06/2025How do language models track mental states of each character in a story, often referred to as Theory of Mind? We reverse-engineered how LLaMA-3-70B-Instruct handles a belief-tracking task and found something surprising: it uses mechanisms strikingly similar to pointer variables in C programming! 25918
Reposted by David AtkinsonDavid Bau @davidbau.bsky.social · 16/03/2025What will be the linchpin for AI dominance? Read our NSF/OSTP recommendations written with Goodfire's Tom McGrath tommcgrath.github.io, Transluce's Sarah Schwettmann cogconfluence.com, MIT's Dylan Hadfield-Menell @dhadfieldmenell.bsky.social TLDR; Dominance comes from **interpretability** 🧵 ↘️ 1218
Reposted by David AtkinsonChantal @chantalsh.bsky.social · 10/03/2025I'm searching for some comp/ling experts to provide a precise definition of “slop” as it refers to text (see: corp.oup.com/word-of-the-...) I put together a google form that should take no longer than 10 minutes to complete: forms.gle/oWxsCScW3dJU... If you can help, I'd appreciate your input! 🙏corp.oup.comOxford Word of the Year 2024 - Oxford University PressThe Oxford Word of the Year 2024 is 'brain rot'. Discover more about the winner, our shortlist, and 20 years of words that reflect the world. 0108
Reposted by David AtkinsonAndrew Lee @ajyl.bsky.social · 20/02/2025Excited about recent reasoning models? What is happening under the hood? Join ARBOR: Analysis of Reasoning Behaviors thru *Open Research* - a radically open collaboration to reverse-engineer reasoning models! Learn more: arborproject.github.io 1/Narborproject.github.ioARBOR 1133
Reposted by David AtkinsonEkdeep Singh @ ICML @ekdeepl.bsky.social · 16/02/2025New paper–accepted as *spotlight* at #ICLR2025! 🧵👇 We show a competition dynamic between several algorithms splits a toy model’s ICL abilities into four broad phases of train/test settings! This means ICL is akin to a mixture of different algorithms, not a monolithic ability. 2325
Reposted by David AtkinsonDavid Bau @davidbau.bsky.social · 07/12/2024PhD Applicants: remember that the Northeastern Computer Science PhD application deadline is Dec 15. It's a terrific time to do a PhD, with so many interesting things happening in AI. Apply here: www.khoury.northeastern.edu/apply/phd-ap...khoury.northeastern.eduPhD Apply - Khoury College of Computer Sciences 0335