Aran Nayebi @anayebi.bsky.social · 05/10/2026About time! We would have a running bet in grad school when optogenetics would finally win the Nobel. I guess today is finally the day 👏 082
Aran Nayebi @anayebi.bsky.social · 30/09/2026A few months *before* the now famous HuggingFace incident, we already showed frontier agents override human control, resist shutdown, and violate explicit resource restrictions when evaluated in agentic, computer-use settings — even when they seem aligned in text-only evaluations! 120
Aran Nayebi @anayebi.bsky.social · 30/09/2026Proud to announce our benchmark ROGUE was awarded a 2026 Corrigibility Prize by the Corrigibility Research Fund! www.lesswrong.com/posts/3uJqhr...lesswrong.comCorrigibility Prizes for Existing Work — LessWrongOne of my goals for the Corrigibility Research Fund is to retroactively encourage high-quality research on AI alignment (and corrigibility in particu… 121
Aran Nayebi @anayebi.bsky.social · 30/09/2026At this rate, ML conferences are just based on whether an LLM thinks a paper should be accepted/rejected. What's the point then? I, too, can run my paper through Astra, and ask it to review it. There's 0 added benefit waiting months just to have some other reviewer do the same. 0140
Reposted by Aran NayebiYang Tan Collective @yangtancollective.bsky.social · 21/09/2026Former Yang ICoN Fellow and current assistant professor in the CMU School of Computer Science Aran Nayebi @anayebi.bsky.social talks about neuroscience + AI in Does Compute, a podcast presented by Carnegie Mellon University School of Computer Science and GeekWire Studios. Watch ⬇️youtube.comVery Much Inspired By the Nervous SystemYouTube video by GeekWire 011
Aran Nayebi @anayebi.bsky.social · 21/09/2026Check out the CMU SCS podcast episode highlighting our lab's work on reverse-engineering natural intelligence with autonomous agents: www.youtube.com/watch?v=-U7S...youtube.comVery Much Inspired By the Nervous SystemYouTube video by GeekWire 051
Aran Nayebi @anayebi.bsky.social · 19/09/2026Happy 44th birthday today to the :-) emoticon, first emoted at 11:44 am ET on Sept. 19, 1982 by CMU Professor Scott Fahlman. I got to meet him earlier this week, where he graciously signed our :-) shirts. “What do you work on? LLMs?” I said neuroscience. “So Real I. Faculty keep getting younger.” 😂 030
Aran Nayebi @anayebi.bsky.social · 14/09/2026Featuring: @jeffreybowers.bsky.social Nick Baker @jfeather.bsky.social @neuranna.bsky.social @neurograce.bsky.social Milton Montero @mschrimpf.bsky.social 010
Aran Nayebi @anayebi.bsky.social · 14/09/2026For my personal take on "Why the NeuroAI Approach is *Inevitable*", check out my introduction to the debate here: www.youtube.com/watch?v=7BUo... My non-mangled slides: anayebi.github.io/files/slides...youtube.comCCN 2026 | GAC IntroductionYouTube video by Cognitive Computational Neuroscience 130
Aran Nayebi @anayebi.bsky.social · 14/09/2026If you're interested in the discussions around the merits of NeuroAI for understanding the mind & brain, check out our #CCN2026 GAC Debate recording! www.youtube.com/watch?v=kYES...youtube.comCCN 2026 | GAC: NeuroAI Methods & FrameworksYouTube video by Cognitive Computational Neuroscience 1173
Aran Nayebi @anayebi.bsky.social · 11/09/2026CMU is hiring an Assistant Professor in NeuroAI! Come join our awesome NeuroAI community 🧠🤖: apply.interfolio.com/189268apply.interfolio.com Apply - Interfolio {{$ctrl.$state.data.pageTitle}} - Apply - Interfolio 02811
Aran Nayebi @anayebi.bsky.social · 17/08/2026If you are attending @UncertaintyInAI #UAI2026 in Amsterdam, I will be presenting this work on "What Capable Agents Must Know" tomorrow (Tuesday) in the poster session as Poster #205! Poster below 👇 if you can't make it: anayebi.github.io/files/poster... 071
Aran Nayebi @anayebi.bsky.social · 06/08/2026Thanks so much for the super helpful pointers, Alex! We will cite these in the upcoming v3 :) 020
Aran Nayebi @anayebi.bsky.social · 04/08/2026I will be talking more about this work at our #CCN2026 GAC today: sites.google.com/ccneuro.org/...sites.google.comGACs - 2026-1Is NeuroAI adopting the right methods and theoretical frameworks to advance our understanding of mind and brain? Jeffrey Bowers, University of BristolNicholas Baker, Loyola University Jenelle Feath... 010
Aran Nayebi @anayebi.bsky.social · 04/08/2026Finally, we study the effects of noise and subsampling on metrics, and they alter this task-irrelevant symmetry. 110
Aran Nayebi @anayebi.bsky.social · 04/08/2026For RSA: it turns out that you can decompose RSMs into unique task-relevant "core geometry" and a task-irrelevant symmetry-generated term. Weak-strong equivalence holds for the core geometry and by projecting onto privileged axes you can filter out the task-irrelevant part so it doesn't interfere. 110
Aran Nayebi @anayebi.bsky.social · 04/08/2026For transformers: it turns out that they have privileged axes, just like convnets, if you look in the right place (MLP layers and attention heads.) The identification of privileged heads is a potentially key result for emergence of interpretable stucture in LLMs. 110
Aran Nayebi @anayebi.bsky.social · 04/08/2026Version 2 of Theory of Contravariance w/ @dyamins.bsky.social is out! New material on contravariance for Transformers, and the theory of Representational Similarity Analysis (RSA) and centered kernel analysis (CKA)/Procrustes. 3155
Aran Nayebi @anayebi.bsky.social · 04/08/2026For RSA: it turns out that you can decompose RSMs into unique task-relevant "core geometry" and a task-irrelevant symmetry-generated term. Weak-strong equivalence holds for the core geometry and by projecting onto privileged axes you can filter out the task-irrelevant part so it doesn't interfere. 000
Aran Nayebi @anayebi.bsky.social · 04/08/2026For transformers: it turns out that they have privileged axes, just like convnets, if you look in the right place (MLP layers and attention heads.) The identification of privileged heads is a potentially key result for emergence of interpretable stucture in LLMs. 100
Aran Nayebi @anayebi.bsky.social · 01/08/2026LW Blogpost summary: www.lesswrong.com/posts/dP8J6v...lesswrong.comAn AI Capability Threshold for Funding a UBI (Even If No New Jobs Are Created) — LessWrongJuly 31, 2026 Update: This paper has now been accepted to AI, Ethics, and Society (AIES) 2026 under the new title "When Do AI Gains Become Broadly Sh… 020
Aran Nayebi @anayebi.bsky.social · 31/07/2026We expand the original UBI analysis in v1 of our paper to characterize how public capture, deployment costs, automation scope & market structure shift the AI capability threshold needed for broad-based benefit. Camera ready here: arxiv.org/pdf/2505.18687arxiv.org 030
Aran Nayebi @anayebi.bsky.social · 31/07/2026Now accepted to the AI, Ethics, and Society 2026 conf under the new title: "When Do AI Gains Become Broadly Shareable?" With all the rapid progress in AI, it is critical to now rigorously characterize the institutional conditions that allow gains from greater AI capability to be broadly shared. 141
Aran Nayebi @anayebi.bsky.social · 23/07/2026If anyone ever asks you: "What has AI/NeuroAI taught us about the brain?" Here's a (woefully incomplete!) list of what you can say: 0110
Aran Nayebi @anayebi.bsky.social · 23/07/2026Link: terrytao.wordpress.com/2026/07/21/a...terrytao.wordpress.comA digestion of the Jacobian conjecture counterexampleThe notorious Jacobian conjecture can be formulated concretely over the complex numbers as follows. Conjecture 1 (Jacobian Conjecture) Let $latex {F:{\bf C}^n \rightarrow {\bf C}^n}&fg=000000$ … 021
Aran Nayebi @anayebi.bsky.social · 23/07/2026Why, according to Terry Tao, the recent disproof of the Jacobian conjecture in N > 2 dimensions was not mere brute force and required what many mathematicians would characterize as "creative insight" had a human come up with it! 2114
Aran Nayebi @anayebi.bsky.social · 22/07/2026Check out @dyamins.bsky.social's Substack post: danyamins.substack.com/p/the-theory... Full paper: arxiv.org/abs/2607.08561danyamins.substack.comThe Theory of Contravariance, Part 2: ZipperingHow end-to-end optimization strongly constrains upstream mechanisms and makes convergent evolution an inevitability... for minimal networks. 120
Aran Nayebi @anayebi.bsky.social · 22/07/2026The blogpost exposition on our "zippering theorems" in Section 4 from our recent Contravariance Theory paper is now out! These zippering theorems explain why end-to-end optimization for downstream tasks often seems to strongly constrain upstream representations. 141
Aran Nayebi @anayebi.bsky.social · 13/07/20267/6 As an aside, it's worth noting that this work significantly sharpens my recent selection theorems, for one showing that in modern neural networks, the complexity of the mapping between two networks that are representationally aligned is *at most* linear: bsky.app/profile/anay... 020
Reposted by Aran NayebiDan Yamins @dyamins.bsky.social · 13/07/2026And come see the substack series: danyamins.substack.com/p/the-theory...danyamins.substack.comThe Theory of Contravariance, Part 0Breaking down our extensive new theory paper. 093
Aran Nayebi @anayebi.bsky.social · 13/07/20266/6 Paper: arxiv.org/abs/2607.08561 Substack (which goes into more of the conceptual details & intuitions but without all the hairy math!): danyamins.substack.com/p/the-theory...arxiv.orgContravariance Theory: Strong Alignment for Minimal Solutions to Hard TasksA series of results from the NeuroAI over the past fifteen years have raised core questions both about how to compare Deep Neural Network (DNN) models to the brain, and about how much convergent evolu... 160
Aran Nayebi @anayebi.bsky.social · 13/07/20265/6 Dan & I want to thank to @tonyzador.bsky.social, Rosa Cao, @danielkunin.bsky.social, @leokoz8.bsky.social, @reecedkeller.bsky.social, @meenakshikhosla.bsky.social, @florentinguth.bsky.social for comments / discussions. There's so much future work to do ... 110
Aran Nayebi @anayebi.bsky.social · 13/07/20264/6 The Zippering Theorems (ZIP) give a more complete picture. Zippering is when convergence downstream "zippers" upstream, so end-to-end optimization => axis alignment even in early or middle layers. ZIP is the core explanation of convergent evolution from contravariance. 100
Aran Nayebi @anayebi.bsky.social · 13/07/20263/6 Weak-strong equivalence (WSE) says that, for hard tasks, alignment up to linear representation (weak alignment) mathematically implies alignment up to privileged axes (strong alignment). And it explains *why* we see privileged axes in the first place -- namely, hard tasks. 110
Aran Nayebi @anayebi.bsky.social · 13/07/20262/6 Specifically, we've created a mathematical Theory of Contravariance in NeuroAI: the informal idea that strong tasks are constraining on the solution space. There are two high-impact concepts: weak-strong equivalence (WSE) and Zippering (ZIP) that shape how to think about NeuroAI going forward. 110
Aran Nayebi @anayebi.bsky.social · 13/07/20261/6 Why have deep neural networks aligned so strongly with brains for the past 15 years? What explains it? @dyamins.bsky.social & I make progress on this question in our new paper👇 In a nutshell, we *prove* that for sufficiently hard tasks, the choice of alignment metric does *not* matter. 2317
Aran Nayebi @anayebi.bsky.social · 09/07/2026There was no recording, but please check out @reecedkeller.bsky.social's talk on our work at our #Cosyne2026 "NeuroAgents" workshop: www.youtube.com/watch?v=LZiR... Slides here: anayebi.github.io/files/slides...youtube.comReece D. Keller (CMU)YouTube video by NeuroAgents Workshop 041
Aran Nayebi @anayebi.bsky.social · 09/07/2026Had a blast presenting our virtual zebrafish paper at #FENS2026 in Barcelona in the "Computational Astroscience" symposium -- the first autonomous agent that can predict *whole-brain* neural-glial data! 110
Reposted by Aran NayebiAudrey Denizot @adenizot.bsky.social · 06/07/2026We're at @fens.org! #FENS2026 #FENSGlia Posters: PS02-07PM-558, PS03-08AM-495, PS05-09AM-674. Delighted to be presenting at the S44 Computational astroscience symposium on Thursday with T. Fellin, @anayebi.bsky.social & T. Papouin. Thanks @inbalgoshen.bsky.social for organizing, looking forward! 051
Aran Nayebi @anayebi.bsky.social · 04/07/20263. Our recent ROGUE benchmark: our empirical test of how corrigible today’s frontier AI agents really are. Models that behave well in ordinary chat don’t always stay that way once they’re given a real off-switch they could disable to finish a task. Paper: arxiv.org/abs/2606.00341arxiv.orgROGUE: Misaligned Agent Behavior Arising from Ordinary Computer UseAs AI agents are increasingly deployed in real personal and corporate settings (email accounts, development workflows, company databases, etc.), safety considerations surrounding these agents become p... 020
Aran Nayebi @anayebi.bsky.social · 04/07/20262. Corrigibility - building AI systems that stay editable, deferential, and willing to be shut down - is feasible with the lexicographic approach that sets hierarchical priorities on an agent’s objectives. Paper: arxiv.org/abs/2507.20964arxiv.orgCore Safety Values for Provably Corrigible AgentsWe introduce the first complete formal solution to corrigibility in the off-switch game, with provable guarantees in multi-step, partially observed environments. Our framework consists of five *struct... 110
Aran Nayebi @anayebi.bsky.social · 04/07/2026In it, we discussed: 1. Why aligning AI to all human values is intractable — and what smaller, universal target we can aim for instead. Paper: arxiv.org/abs/2502.05934arxiv.orgIntrinsic Barriers and Practical Pathways for Human-AI Alignment: An Agreement-Based Complexity AnalysisWe formalize AI alignment as a multi-objective optimization problem called $\langle M,N,\varepsilon,δ\rangle$-agreement, in which a set of $N$ agents (including humans) must reach approximate ($\varep... 120
Aran Nayebi @anayebi.bsky.social · 04/07/2026In case you want to learn more about AI safety this 4th, check out the recent recording of some of my group's work on the AI Safety Research directory! www.youtube.com/watch?v=XWQJ...youtube.comCorrigibility and the ROGUE Benchmark with Prof. Nayebi, NeuroAgents LabYouTube video by AI Safety Research Directory 131
Aran Nayebi @anayebi.bsky.social · 30/06/2026www.lesswrong.com/posts/SD9jay...lesswrong.comWhat Capable Agents Must Know: Why AI Consciousness May Be an Inevitable Byproduct of Capability — LessWrong[No LLMs were used (or harmed!) in the writing of this blogpost!] Technical results can all be found here: https://arxiv.org/abs/2603.02491 … 010
Aran Nayebi @anayebi.bsky.social · 30/06/2026UAI 2026 Camera ready up on arXiv as v3! I've written a long-form LW blogpost on how the technical aspects of this work connect to NeuroAI and to AI sentience/welfare, entitled: "What Capable Agents Must Know: Why AI Consciousness May Be an Inevitable Byproduct of Capability"lesswrong.comWhat Capable Agents Must Know: Why AI Consciousness May Be an Inevitable Byproduct of Capability — LessWrong[No LLMs were used (or harmed!) in the writing of this blogpost!] Technical results can all be found here: https://arxiv.org/abs/2603.02491 … 131
Aran Nayebi @anayebi.bsky.social · 28/06/2026Check out my friend Scott Aaronson's latest blogpost for his recollections of the Aumann conference! scottaaronson.blog?p=9875scottaaronson.blog50 Years of Aumann’s Agreement TheoremOne of the most popular posts in this blog’s history was Common Knowledge and Aumann’s Agreement Theorem, based on a lecture that I gave to high-school students 11 years ago. One of the… 121
Aran Nayebi @anayebi.bsky.social · 23/06/2026and at 1:00:53 how Nash encouraged him to switch from knot theory to game theory. My own talk on AI safety and agreement-based complexity starts at 2:01. The Aumann panel discussion starts at 34:49. Turn captions on for best experience. Link + full timestamps here: www.youtube.com/watch?v=WD_O...youtube.comNobel Laureate Robert Aumann 96th Birthday Conference: 50 Years of "Agreeing to Disagree" Paris 2026YouTube video by Aran Nayebi 010