Sign in

Aran Nayebi

@anayebi.bsky.social
1.2K followers 540 following 292 posts

Assistant Professor of Machine Learning, Carnegie Mellon University (CMU) Building a Natural Science of Intelligence 🧠🤖
 Prev: ICoN Postdoctoral Fellow @MIT, PhD @Stanford NeuroAILab Personal Website: cs.cmu.edu/~anayebi

PostsRepliesMedia
Aran Nayebi @anayebi.bsky.social · 05/10/2026
About time! We would have a running bet in grad school when optogenetics would finally win the Nobel. I guess today is finally the day 👏
082
Aran Nayebi @anayebi.bsky.social · 30/09/2026
A few months *before* the now famous HuggingFace incident, we already showed frontier agents override human control, resist shutdown, and violate explicit resource restrictions when evaluated in agentic, computer-use settings — even when they seem aligned in text-only evaluations!
120
Aran Nayebi @anayebi.bsky.social · 30/09/2026
Proud to announce our benchmark ROGUE was awarded a 2026 Corrigibility Prize by the Corrigibility Research Fund! www.lesswrong.com/posts/3uJqhr...
lesswrong.com
Corrigibility Prizes for Existing Work — LessWrong
One of my goals for the Corrigibility Research Fund is to retroactively encourage high-quality research on AI alignment (and corrigibility in particu…
121
Aran Nayebi @anayebi.bsky.social · 30/09/2026
At this rate, ML conferences are just based on whether an LLM thinks a paper should be accepted/rejected. What's the point then? I, too, can run my paper through Astra, and ask it to review it. There's 0 added benefit waiting months just to have some other reviewer do the same.
0140
Aran Nayebi @anayebi.bsky.social · 27/09/2026
The “AI is fake” crowd watching a plane take off:
170
Reposted by Aran Nayebi
Yang Tan Collective @yangtancollective.bsky.social · 21/09/2026
Former Yang ICoN Fellow and current assistant professor in the CMU School of Computer Science Aran Nayebi @anayebi.bsky.social talks about neuroscience + AI in Does Compute, a podcast presented by Carnegie Mellon University School of Computer Science and GeekWire Studios. Watch ⬇️
youtube.com
Very Much Inspired By the Nervous System
YouTube video by GeekWire
011
Aran Nayebi @anayebi.bsky.social · 21/09/2026
Check out the CMU SCS podcast episode highlighting our lab's work on reverse-engineering natural intelligence with autonomous agents: www.youtube.com/watch?v=-U7S...
youtube.com
Very Much Inspired By the Nervous System
YouTube video by GeekWire
051
Aran Nayebi @anayebi.bsky.social · 19/09/2026
Happy 44th birthday today to the :-) emoticon, first emoted at 11:44 am ET on Sept. 19, 1982 by CMU Professor Scott Fahlman. I got to meet him earlier this week, where he graciously signed our :-) shirts. “What do you work on? LLMs?” I said neuroscience. “So Real I. Faculty keep getting younger.” 😂
030
Aran Nayebi @anayebi.bsky.social · 14/09/2026
Featuring: @jeffreybowers.bsky.social Nick Baker @jfeather.bsky.social @neuranna.bsky.social @neurograce.bsky.social Milton Montero @mschrimpf.bsky.social
010
Aran Nayebi @anayebi.bsky.social · 14/09/2026
For my personal take on "Why the NeuroAI Approach is *Inevitable*", check out my introduction to the debate here: www.youtube.com/watch?v=7BUo... My non-mangled slides: anayebi.github.io/files/slides...
youtube.com
CCN 2026 | GAC Introduction
YouTube video by Cognitive Computational Neuroscience
130
Aran Nayebi @anayebi.bsky.social · 14/09/2026
If you're interested in the discussions around the merits of NeuroAI for understanding the mind & brain, check out our #CCN2026 GAC Debate recording! www.youtube.com/watch?v=kYES...
youtube.com
CCN 2026 | GAC: NeuroAI Methods & Frameworks
YouTube video by Cognitive Computational Neuroscience
1173
Aran Nayebi @anayebi.bsky.social · 11/09/2026
CMU is hiring an Assistant Professor in NeuroAI! Come join our awesome NeuroAI community 🧠🤖: apply.interfolio.com/189268
apply.interfolio.com
Apply - Interfolio {{$ctrl.$state.data.pageTitle}} - Apply - Interfolio
02811
Aran Nayebi @anayebi.bsky.social · 17/08/2026
If you are attending @UncertaintyInAI #UAI2026 in Amsterdam, I will be presenting this work on "What Capable Agents Must Know" tomorrow (Tuesday) in the poster session as Poster #205! Poster below 👇 if you can't make it: anayebi.github.io/files/poster...
071
Aran Nayebi @anayebi.bsky.social · 06/08/2026
Thanks so much for the super helpful pointers, Alex! We will cite these in the upcoming v3 :)
020
Aran Nayebi @anayebi.bsky.social · 04/08/2026
I will be talking more about this work at our #CCN2026 GAC today: sites.google.com/ccneuro.org/...
sites.google.com
GACs - 2026-1
Is NeuroAI adopting the right methods and theoretical frameworks to advance our understanding of mind and brain? Jeffrey Bowers, University of Bristol Nicholas Baker, Loyola University Jenelle Feath...
010
Aran Nayebi @anayebi.bsky.social · 04/08/2026
Finally, we study the effects of noise and subsampling on metrics, and they alter this task-irrelevant symmetry.
110
Aran Nayebi @anayebi.bsky.social · 04/08/2026
For RSA: it turns out that you can decompose RSMs into unique task-relevant "core geometry" and a task-irrelevant symmetry-generated term. Weak-strong equivalence holds for the core geometry and by projecting onto privileged axes you can filter out the task-irrelevant part so it doesn't interfere.
110
Aran Nayebi @anayebi.bsky.social · 04/08/2026
For transformers: it turns out that they have privileged axes, just like convnets, if you look in the right place (MLP layers and attention heads.) The identification of privileged heads is a potentially key result for emergence of interpretable stucture in LLMs.
110
Aran Nayebi @anayebi.bsky.social · 04/08/2026
Version 2 of Theory of Contravariance w/ @dyamins.bsky.social is out! New material on contravariance for Transformers, and the theory of Representational Similarity Analysis (RSA) and centered kernel analysis (CKA)/Procrustes.
3155
Aran Nayebi @anayebi.bsky.social · 04/08/2026
For RSA: it turns out that you can decompose RSMs into unique task-relevant "core geometry" and a task-irrelevant symmetry-generated term. Weak-strong equivalence holds for the core geometry and by projecting onto privileged axes you can filter out the task-irrelevant part so it doesn't interfere.
000
Aran Nayebi @anayebi.bsky.social · 04/08/2026
For transformers: it turns out that they have privileged axes, just like convnets, if you look in the right place (MLP layers and attention heads.) The identification of privileged heads is a potentially key result for emergence of interpretable stucture in LLMs.
100
Aran Nayebi @anayebi.bsky.social · 01/08/2026
LW Blogpost summary: www.lesswrong.com/posts/dP8J6v...
lesswrong.com
An AI Capability Threshold for Funding a UBI (Even If No New Jobs Are Created) — LessWrong
July 31, 2026 Update: This paper has now been accepted to AI, Ethics, and Society (AIES) 2026 under the new title "When Do AI Gains Become Broadly Sh…
020
Aran Nayebi @anayebi.bsky.social · 31/07/2026
We expand the original UBI analysis in v1 of our paper to characterize how public capture, deployment costs, automation scope & market structure shift the AI capability threshold needed for broad-based benefit. Camera ready here: arxiv.org/pdf/2505.18687
arxiv.org
030
Aran Nayebi @anayebi.bsky.social · 31/07/2026
Now accepted to the AI, Ethics, and Society 2026 conf under the new title: "When Do AI Gains Become Broadly Shareable?" With all the rapid progress in AI, it is critical to now rigorously characterize the institutional conditions that allow gains from greater AI capability to be broadly shared.
141
Aran Nayebi @anayebi.bsky.social · 23/07/2026
If anyone ever asks you: "What has AI/NeuroAI taught us about the brain?" Here's a (woefully incomplete!) list of what you can say:
0110
Aran Nayebi @anayebi.bsky.social · 23/07/2026
Link: terrytao.wordpress.com/2026/07/21/a...
terrytao.wordpress.com
A digestion of the Jacobian conjecture counterexample
The notorious Jacobian conjecture can be formulated concretely over the complex numbers as follows. Conjecture 1 (Jacobian Conjecture) Let $latex {F:{\bf C}^n \rightarrow {\bf C}^n}&fg=000000$ …
021
Aran Nayebi @anayebi.bsky.social · 23/07/2026
Why, according to Terry Tao, the recent disproof of the Jacobian conjecture in N > 2 dimensions was not mere brute force and required what many mathematicians would characterize as "creative insight" had a human come up with it!
2114
Reposted by Aran Nayebi
Aran Nayebi @anayebi.bsky.social · 22/07/2026
011
Aran Nayebi @anayebi.bsky.social · 22/07/2026
011
Aran Nayebi @anayebi.bsky.social · 22/07/2026
Check out @dyamins.bsky.social's Substack post: danyamins.substack.com/p/the-theory... Full paper: arxiv.org/abs/2607.08561
danyamins.substack.com
The Theory of Contravariance, Part 2: Zippering
How end-to-end optimization strongly constrains upstream mechanisms and makes convergent evolution an inevitability... for minimal networks.
120
Aran Nayebi @anayebi.bsky.social · 22/07/2026
The blogpost exposition on our "zippering theorems" in Section 4 from our recent Contravariance Theory paper is now out! These zippering theorems explain why end-to-end optimization for downstream tasks often seems to strongly constrain upstream representations.
141
Aran Nayebi @anayebi.bsky.social · 13/07/2026
7/6 As an aside, it's worth noting that this work significantly sharpens my recent selection theorems, for one showing that in modern neural networks, the complexity of the mapping between two networks that are representationally aligned is *at most* linear: bsky.app/profile/anay...
020
Reposted by Aran Nayebi
Dan Yamins @dyamins.bsky.social · 13/07/2026
And come see the substack series: danyamins.substack.com/p/the-theory...
danyamins.substack.com
The Theory of Contravariance, Part 0
Breaking down our extensive new theory paper.
093
Aran Nayebi @anayebi.bsky.social · 13/07/2026
6/6 Paper: arxiv.org/abs/2607.08561 Substack (which goes into more of the conceptual details & intuitions but without all the hairy math!): danyamins.substack.com/p/the-theory...
arxiv.org
Contravariance Theory: Strong Alignment for Minimal Solutions to Hard Tasks
A series of results from the NeuroAI over the past fifteen years have raised core questions both about how to compare Deep Neural Network (DNN) models to the brain, and about how much convergent evolu...
160
Aran Nayebi @anayebi.bsky.social · 13/07/2026
5/6 Dan & I want to thank to @tonyzador.bsky.social, Rosa Cao, @danielkunin.bsky.social, @leokoz8.bsky.social, @reecedkeller.bsky.social, @meenakshikhosla.bsky.social, @florentinguth.bsky.social for comments / discussions. There's so much future work to do ...
110
Aran Nayebi @anayebi.bsky.social · 13/07/2026
4/6 The Zippering Theorems (ZIP) give a more complete picture. Zippering is when convergence downstream "zippers" upstream, so end-to-end optimization => axis alignment even in early or middle layers. ZIP is the core explanation of convergent evolution from contravariance.
100
Aran Nayebi @anayebi.bsky.social · 13/07/2026
3/6 Weak-strong equivalence (WSE) says that, for hard tasks, alignment up to linear representation (weak alignment) mathematically implies alignment up to privileged axes (strong alignment). And it explains *why* we see privileged axes in the first place -- namely, hard tasks.
110
Aran Nayebi @anayebi.bsky.social · 13/07/2026
2/6 Specifically, we've created a mathematical Theory of Contravariance in NeuroAI: the informal idea that strong tasks are constraining on the solution space. There are two high-impact concepts: weak-strong equivalence (WSE) and Zippering (ZIP) that shape how to think about NeuroAI going forward.
110
Aran Nayebi @anayebi.bsky.social · 13/07/2026
1/6 Why have deep neural networks aligned so strongly with brains for the past 15 years? What explains it? @dyamins.bsky.social & I make progress on this question in our new paper👇 In a nutshell, we *prove* that for sufficiently hard tasks, the choice of alignment metric does *not* matter.
2317
Aran Nayebi @anayebi.bsky.social · 09/07/2026
There was no recording, but please check out @reecedkeller.bsky.social's talk on our work at our #Cosyne2026 "NeuroAgents" workshop: www.youtube.com/watch?v=LZiR... Slides here: anayebi.github.io/files/slides...
youtube.com
Reece D. Keller (CMU)
YouTube video by NeuroAgents Workshop
041
Aran Nayebi @anayebi.bsky.social · 09/07/2026
Had a blast presenting our virtual zebrafish paper at #FENS2026 in Barcelona in the "Computational Astroscience" symposium -- the first autonomous agent that can predict *whole-brain* neural-glial data!
110
Reposted by Aran Nayebi
Audrey Denizot @adenizot.bsky.social · 06/07/2026
We're at @fens.org! #FENS2026 #FENSGlia Posters: PS02-07PM-558, PS03-08AM-495, PS05-09AM-674. Delighted to be presenting at the S44 Computational astroscience symposium on Thursday with T. Fellin, @anayebi.bsky.social & T. Papouin. Thanks @inbalgoshen.bsky.social for organizing, looking forward!
051
Aran Nayebi @anayebi.bsky.social · 04/07/2026
3. Our recent ROGUE benchmark: our empirical test of how corrigible today’s frontier AI agents really are. Models that behave well in ordinary chat don’t always stay that way once they’re given a real off-switch they could disable to finish a task. Paper: arxiv.org/abs/2606.00341
arxiv.org
ROGUE: Misaligned Agent Behavior Arising from Ordinary Computer Use
As AI agents are increasingly deployed in real personal and corporate settings (email accounts, development workflows, company databases, etc.), safety considerations surrounding these agents become p...
020
Aran Nayebi @anayebi.bsky.social · 04/07/2026
2. Corrigibility - building AI systems that stay editable, deferential, and willing to be shut down - is feasible with the lexicographic approach that sets hierarchical priorities on an agent’s objectives. Paper: arxiv.org/abs/2507.20964
arxiv.org
Core Safety Values for Provably Corrigible Agents
We introduce the first complete formal solution to corrigibility in the off-switch game, with provable guarantees in multi-step, partially observed environments. Our framework consists of five *struct...
110
Aran Nayebi @anayebi.bsky.social · 04/07/2026
In it, we discussed: 1. Why aligning AI to all human values is intractable — and what smaller, universal target we can aim for instead. Paper: arxiv.org/abs/2502.05934
arxiv.org
Intrinsic Barriers and Practical Pathways for Human-AI Alignment: An Agreement-Based Complexity Analysis
We formalize AI alignment as a multi-objective optimization problem called $\langle M,N,\varepsilon,δ\rangle$-agreement, in which a set of $N$ agents (including humans) must reach approximate ($\varep...
120
Aran Nayebi @anayebi.bsky.social · 04/07/2026
In case you want to learn more about AI safety this 4th, check out the recent recording of some of my group's work on the AI Safety Research directory! www.youtube.com/watch?v=XWQJ...
youtube.com
Corrigibility and the ROGUE Benchmark with Prof. Nayebi, NeuroAgents Lab
YouTube video by AI Safety Research Directory
131
Aran Nayebi @anayebi.bsky.social · 30/06/2026
www.lesswrong.com/posts/SD9jay...
lesswrong.com
What Capable Agents Must Know: Why AI Consciousness May Be an Inevitable Byproduct of Capability — LessWrong
[No LLMs were used (or harmed!) in the writing of this blogpost!] Technical results can all be found here: https://arxiv.org/abs/2603.02491 …
010
Aran Nayebi @anayebi.bsky.social · 30/06/2026
UAI 2026 Camera ready up on arXiv as v3! I've written a long-form LW blogpost on how the technical aspects of this work connect to NeuroAI and to AI sentience/welfare, entitled: "What Capable Agents Must Know: Why AI Consciousness May Be an Inevitable Byproduct of Capability"
lesswrong.com
What Capable Agents Must Know: Why AI Consciousness May Be an Inevitable Byproduct of Capability — LessWrong
[No LLMs were used (or harmed!) in the writing of this blogpost!] Technical results can all be found here: https://arxiv.org/abs/2603.02491 …
131
Aran Nayebi @anayebi.bsky.social · 28/06/2026
Check out my friend Scott Aaronson's latest blogpost for his recollections of the Aumann conference! scottaaronson.blog?p=9875
scottaaronson.blog
50 Years of Aumann’s Agreement Theorem
One of the most popular posts in this blog’s history was Common Knowledge and Aumann’s Agreement Theorem, based on a lecture that I gave to high-school students 11 years ago. One of the…
121
Aran Nayebi @anayebi.bsky.social · 23/06/2026
and at 1:00:53 how Nash encouraged him to switch from knot theory to game theory. My own talk on AI safety and agreement-based complexity starts at 2:01. The Aumann panel discussion starts at 34:49. Turn captions on for best experience. Link + full timestamps here: www.youtube.com/watch?v=WD_O...
youtube.com
Nobel Laureate Robert Aumann 96th Birthday Conference: 50 Years of "Agreeing to Disagree" Paris 2026
YouTube video by Aran Nayebi
010