Sign in

Steve Byrnes

@stevebyrnes.bsky.social
3K followers 88 following 427 posts

Researching Artificial General Intelligence Safety, via thinking about neuroscience and algorithms, at Astera Institute. sjbyrnes.com/agi.html

PostsRepliesMedia
Steve Byrnes @stevebyrnes.bsky.social · 18/09/2026
Blog post: “Pretraining data, not verifiability, is why LLMs are especially good at math (and coding)” www.lesswrong.com/posts/xvdngZ...
lesswrong.com
Pretraining data, not verifiability, is why LLMs are especially good at math (and coding) — LessWrong
A common take I’ve been hearing is: “LLMs are especially good at math because math is easy to verify”. But that story doesn’t make much sense.…
061
Steve Byrnes @stevebyrnes.bsky.social · 29/08/2026
Blog post: “Tales of rebellion against externally-opaque meritocracies” www.lesswrong.com/posts/m8cP9K...
lesswrong.com
010
Steve Byrnes @stevebyrnes.bsky.social · 25/08/2026
I strongly agree with Cal-Newport-when-talking-about-mundane-LLM-risks, that LLMs are not synonymous with AI. AI is a whole field, including yet-to-be-invented paradigms wildly different from LLMs. He should go tell this important news to his friend Cal-Newport-when-talking-about-superintelligence.
020
Steve Byrnes @stevebyrnes.bsky.social · 10/08/2026
Blog post: “Four LLM loss functions → four flavors of LLM misalignment” www.alignmentforum.org/posts/GRmvZs...
000
Steve Byrnes @stevebyrnes.bsky.social · 08/08/2026
“AI companies seem to love making this story about the AI being scary, because if you blame the AI, you aren’t blaming them.” (screenshot of a blog post)Images from Terminator 2: Sarah Connor: “This is all your fault! Die you motherf***r!” Miles Dyson: “But it can’t be my fault, because it’s Skynet’s fault” Sarah Connor: “Oh! That makes perfect sense! Sorry, my bad! Carry on!”
010
Steve Byrnes @stevebyrnes.bsky.social · 07/08/2026
It’s funny because these are all based on a 1967 experiment that had already failed to replicate in 1969. (& failed again in 2003.) It’s also funny because it’s pitched as a fun test that a group can do in 2 minutes at a party, but I guess none of these writers actually tried before publishing?
030
Steve Byrnes @stevebyrnes.bsky.social · 27/07/2026
Blog post: “RL & search is a terrifying way to build AGI (an FAQ)” www.alignmentforum.org/posts/KHyBoc... The FAQ has all your burning Q’s: Q1: What are you saying? Q2: So you’re saying, don’t build AGI based on RL and/or search & planning? Q3: Why do you think it’s terrifying? 🧵 (1/7)
alignmentforum.org
RL & search is a terrifying way to build AGI (an FAQ) — AI Alignment Forum
Q1: WHAT ARE YOU SAYING? A: My claim here is that if you build artificial general intelligence (AGI) via any algorithm that’s choosing actions via RL and/or model-based search and planning—a giant chu...
230
Steve Byrnes @stevebyrnes.bsky.social · 24/07/2026
Blog post: “LLMs are (still) mostly powered by imitative learning, not RL” www.lesswrong.com/posts/wYpjXR...
lesswrong.com
LLMs are (still) mostly powered by imitative learning, not RL — LessWrong
LLMs get their impressive capabilities from a combination of: (1) Imitative learning from pretraining and SFT data (see my earlier discussion of “LLM pretraining magically transmutes observations into...
000
Steve Byrnes @stevebyrnes.bsky.social · 22/07/2026
Blog post: “Will almost all future companies eventually be founded and run by autonomous AIs?” I think this question is a great conversation-starter for people talking past each other on the future of AI. I go through some common responses, and my own replies (1/3) www.lesswrong.com/posts/BHEcss...
lesswrong.com
Will almost all future companies eventually be founded and run by autonomous AIs? — LessWrong
My belief is that keeping AI under human control would be an unprecedented global challenge, if it’s even possible at all. Many other people believe that AIs will naturally remain as tools under human...
220
Steve Byrnes @stevebyrnes.bsky.social · 20/07/2026
Blog post: What do I mean by “Artificial General Intelligence”? I paint a brief picture of what I’m talking about when I talk about “AGI”. Should seem obvious to many, and obviously wrong to many others! Hope it can help surface disagreements. www.lesswrong.com/posts/nQH2Gh...
lesswrong.com
What do I mean by “Artificial General Intelligence”? — LessWrong
In this post, intended for a broad audience, I will paint a brief picture of what I’m talking about when I talk about “AGI”. It will seem obvious to many people, and obviously wrong to many others! So...
020
Steve Byrnes @stevebyrnes.bsky.social · 13/07/2026
Just updated the “Intro to Brain-Like-AGI Safety” PDF to version 4. Links in QT ↓ A few highlights from the changelog are listed here: www.lesswrong.com/posts/btHmC8...
010
Steve Byrnes @stevebyrnes.bsky.social · 08/07/2026
Blog post: “Notes on technical alignment via human-like social drives” www.alignmentforum.org/posts/rKdS7i...
alignmentforum.org
Notes on technical alignment via human-like social drives — AI Alignment Forum
As my regular readers know (see Intro to Brain-Like-AGI Safety), I’m working on the technical alignment problem for a hypothetical future “brain-like AGI”, with a particular focus on how human social ...
060
Steve Byrnes @stevebyrnes.bsky.social · 01/07/2026
my deranged metascience idea is to create a next-gen crowdsourced ranking of good science papers / books / etc., but we’ll call it “the smelly fart index”, so nobody wants to brag about it on their CV, and therefore nobody bothers to game it
010
Steve Byrnes @stevebyrnes.bsky.social · 21/06/2026
If someone is on a hunger strike for freedom, they’ll say that their desire to eat comes from their innate drives, whereas their desire to fight for freedom comes from “reason” etc. I think this intuition is very wrong; here’s my latest attempt to articulate why www.lesswrong.com/posts/qNZSBq...
120
Steve Byrnes @stevebyrnes.bsky.social · 21/06/2026
I’ll probably annoy all sides with this not-really-an-apology that I just tacked onto the beginning of my LLM-skepticism-related post from Apr 2023
020
Steve Byrnes @stevebyrnes.bsky.social · 12/06/2026
Blog post: “Sympathy for both sides of the egregious misalignment debate” www.alignmentforum.org/posts/DZaZ3f...
alignmentforum.org
Sympathy for both sides of the egregious misalignment debate — AI Alignment Forum
On one side of this debate is Yudkowsky & Soares, who think that (if AI progress continues) we’re on a direct path to egregiously-misaligned, schemin…
010
Steve Byrnes @stevebyrnes.bsky.social · 27/05/2026
“Neuroscience of Human Social Instincts: A Sketch” (2024) (IMO my most important pure neuro work to date) is now version 3 & better than ever! Confused & muddled descriptions are out, streamlined explanations & better terminology are in! Abstract & Changelog screenshots, plus links ↓ ↓
030
Steve Byrnes @stevebyrnes.bsky.social · 11/05/2026
Blog post: “Empowerment, corrigibility, etc. are simple abstractions (of a messed-up ontology)” www.alignmentforum.org/posts/vzHtHH...
000
Steve Byrnes @stevebyrnes.bsky.social · 04/05/2026
I revised 2 old posts based on deeper appreciation for the Orienting Reflex as exemplifying a broader neuroscience motif ↓
110
Steve Byrnes @stevebyrnes.bsky.social · 10/04/2026
Blog post: “Some takes on UV & cancer” • Part 1: In which I use my optical physics background to share some hopefully-uncontroversial observations • Part 2: In which I boldly defy Public Health Orthodoxy on the whole UV situation www.lesswrong.com/posts/t7GeZn...
lesswrong.com
Some takes on UV & cancer — LessWrong
ToC: Part 1: In which I use my optical physics background to share some hopefully-uncontroversial observations. Part 2: In which I boldly defy Public Health Orthodoxy on the whole UV situation
010
Steve Byrnes @stevebyrnes.bsky.social · 18/03/2026
New blog post: “‘Act-based approval-directed agents’, for IDA skeptics” www.alignmentforum.org/posts/RKtTi8...
020
Steve Byrnes @stevebyrnes.bsky.social · 17/03/2026
New blog post: “You can’t imitation-learn how to continual-learn” www.lesswrong.com/posts/9rCTjb...
screenshot of the title and first couple paragraphs of the article at the link
130
Steve Byrnes @stevebyrnes.bsky.social · 17/03/2026
I linkposted a podcast with some excerpts. Title: “Jeremy Howard is bearish on LLMs” [I mean “bearish” compared to most people in my professional circles] www.lesswrong.com/posts/hvun2m...
lesswrong.com
Podcast: Jeremy Howard is bearish on LLMs — LessWrong
Jeremy Howard was recently[1] interviewed on the Machine Learning Street Talk podcast: YouTube link, interactive transcript, PDF transcript. …
010
Steve Byrnes @stevebyrnes.bsky.social · 18/02/2026
New post: “Why we should expect ruthless sociopath ASI” A rift between super-pessimists like me, vs the “merely” AI-concerned, is an intuition that future AI will be kinda like a ruthless sociopath. I claim it’s a sound intuition, but it might seem to come from left field… 1/2
4120
Steve Byrnes @stevebyrnes.bsky.social · 17/02/2026
Blog post: “The brain is a machine that runs an algorithm” www.lesswrong.com/posts/eKGjwR...
lesswrong.com
The brain is a machine that runs an algorithm — LessWrong
Some people say “the brain is a computer”. Other people say “well, the brain is not really a computer, because, like, what’s the hardware versus the software?” I agree: “the brain is a computer” is ki...
1152
Steve Byrnes @stevebyrnes.bsky.social · 06/02/2026
Blog post: “In (highly contingent!) defense of interpretability-in-the-loop ML training”. Using interpretability as input into a loss function / reward function has a bad rap (and deservedly so). But there’s a specific version of it that might work. www.alignmentforum.org/posts/ArXAyz...
alignmentforum.org
000
Steve Byrnes @stevebyrnes.bsky.social · 06/02/2026
Blog post: “The nature of LLM algorithmic progress” bit of Cunningham’s Law energy with this one: spicy hot takes, far outside my area of expertise. Feedback welcome! www.lesswrong.com/posts/sGNFtW...
130
Steve Byrnes @stevebyrnes.bsky.social · 02/02/2026
New blog post: “Are there lessons from high-reliability engineering for AGI safety?” People sometimes suggest that high-reliability engineering is a model for how AGI safety could or should work. I agree in some ways and disagree in other ways. (1/3) www.alignmentforum.org/posts/hiigux...
alignmentforum.org
120
Steve Byrnes @stevebyrnes.bsky.social · 23/01/2026
New version of “Intro to Brain-Like-AGI Safety” is out! Same links as before: • Blog version: www.alignmentforum.org/s/HzcM2dkCq7... • PDF version (v3): osf.io/preprints/os... • Summary video: youtu.be/IXi96sRMKUI • Summary thread: bsky.app/profile/stev... More in thread… 🧵
150
Steve Byrnes @stevebyrnes.bsky.social · 05/01/2026
For anyone who read my “Intuitive Self-Models” series (2024), FYI: I just went through and globally replaced the term “homunculus” with “Active Self”. I’ve come to think that “Active Self” is just a much better term for the specific thing I was trying to talk about. www.lesswrong.com/s/qhdHbCJ3PY...
lesswrong.com
Intuitive Self-Models — LessWrong
This is a rather ambitious series of blog posts, in that I’ll attempt to explain what’s the deal with consciousness, free will, hypnotism, enlightenment, hallucinations, flow states, dissociation, akr...
120
Steve Byrnes @stevebyrnes.bsky.social · 16/12/2025
For archival, citing, & printing purposes, I cross-posted “Neuroscience of human social instincts: a sketch” as a PDF: doi.org/10.5281/zeno... 1 year later, still my proudest & IMO most important pure neuroscience work :) (Blog version is still at: www.lesswrong.com/posts/kYvbHC... )
010
Steve Byrnes @stevebyrnes.bsky.social · 15/12/2025
000
Steve Byrnes @stevebyrnes.bsky.social · 11/12/2025
Blog post: “My AGI safety research—2025 review, ’26 plans”. Been a busy and productive year, progress on many fronts, but I still got my work cut out going forward! www.alignmentforum.org/posts/CF4Z9m...
alignmentforum.org
020
Steve Byrnes @stevebyrnes.bsky.social · 08/12/2025
Blog post: “Reward Function Design: a starter pack”. RL reward functions tend to make ruthless optimizers, when they work at all. But some don’t, and over the past years, I’ve been puzzling about over how. I’ve wound up with a bunch of useful frames and concepts. Here are 5: 🧵
160
Steve Byrnes @stevebyrnes.bsky.social · 08/12/2025
A brief call-to-action blog post: We [desperately] need a field of “Reward Function Design” (1/7) www.alignmentforum.org/posts/oxvnRE...
130
Steve Byrnes @stevebyrnes.bsky.social · 03/12/2025
Blog post: 6 reasons why “alignment-is-hard” discourse seems alien to human intuitions, and vice-versa www.alignmentforum.org/posts/d4HNRd...
screenshot of text from the tl;dr at the top of https://www.alignmentforum.org/posts/d4HNRdw6z7Xqbnu5E/6-reasons-why-alignment-is-hard-discourse-seems-alien-toscreenshot of text from the tl;dr at the top of https://www.alignmentforum.org/posts/d4HNRdw6z7Xqbnu5E/6-reasons-why-alignment-is-hard-discourse-seems-alien-toSimpsons screenshot from https://www.alignmentforum.org/posts/d4HNRdw6z7Xqbnu5E/6-reasons-why-alignment-is-hard-discourse-seems-alien-to . Lisa is saying to Homer: “I mean, if you’re the police, then who will police the police?”
010
Steve Byrnes @stevebyrnes.bsky.social · 12/11/2025
New blog post! “Social drives 2: ‘Approval Reward’, from norm-enforcement to status-seeking”. I try to explain the path from an RL reward function in the brain, to deep truths about the human psyche… www.lesswrong.com/posts/fPxgFH... (1/6)
lesswrong.com
Social drives 2: “Approval Reward”, from norm-enforcement to status-seeking — LessWrong
…Approval Reward is a brain signal that leads to: • Pleasure (positive reward) when my friends and idols seem to have positive feelings about me, or about something related to me, or about what I’m do...
110
Steve Byrnes @stevebyrnes.bsky.social · 10/11/2025
New blog post! “Social drives 1: ‘Sympathy Reward’, from compassion to dehumanization”. This is the 1st of 2 posts building an ever-better bridge that connects from neuroscience & algorithms on one shore, to everyday human experience on the other… www.lesswrong.com/posts/KuBiv9... (1/5)
120
Steve Byrnes @stevebyrnes.bsky.social · 07/10/2025
Blog post: “Excerpts from my neuroscience to-do list” www.lesswrong.com/posts/c6Job6...
lesswrong.com
120
Steve Byrnes @stevebyrnes.bsky.social · 28/09/2025
“The human niche” includes living on every continent, walking on the moon, inventing computers & nuclear weapons, and unraveling the secrets of the universe. We need a term for AI systems that can occupy this “niche”. (1/6)
120
Steve Byrnes @stevebyrnes.bsky.social · 18/09/2025
Just read the new book “If Anyone Builds It, Everyone Dies”. Upshot: Recommended! I ~90% agree with it. Thread: ifanyonebuildsit.com
ifanyonebuildsit.com
If Anyone Builds It, Everyone Dies
The race to superhuman AI risks extinction, but it's not too late to change course.
110
Steve Byrnes @stevebyrnes.bsky.social · 12/09/2025
Blog post: Optical rectennas are not a promising clean energy technology www.lesswrong.com/posts/gKCavz...
lesswrong.com
Optical rectennas are not a promising clean energy technology — LessWrong
“Optical rectennas” (or sometimes “nantennas”) are a technology that is sometimes advertised as a path towards converting solar energy to electricity with higher efficiency than normal solar cells...
261
Steve Byrnes @stevebyrnes.bsky.social · 31/08/2025
Clarification: When I shared this meme 2 years ago, I was referring specifically to traditional task-based fMRI studies. “Functional Connectomics” fMRI studies, by contrast, would be flying overhead in a helicopter, strafing the water with a machine gun
030
Steve Byrnes @stevebyrnes.bsky.social · 25/08/2025
Blog post: “Neuroscience of human sexual attraction triggers (3 hypotheses)” www.lesswrong.com/posts/ktydLo...
010
Steve Byrnes @stevebyrnes.bsky.social · 21/08/2025
Blog post: “Four ways learning Econ makes people dumber re: future AI” www.alignmentforum.org/posts/xJWBof...
alignmentforum.org
Four ways learning Econ makes people dumber re: future AI — AI Alignment Forum
There’s a funny thing where economics education paradoxically makes people DUMBER at thinking about future AI. Econ textbooks teach concepts & frames that are great for most things, but counterproduct...
020
Steve Byrnes @stevebyrnes.bsky.social · 12/08/2025
Uploaded a new PDF version of ↓, with various minor changes accumulated over the last 5 months—a few new paragraphs, new references, typo fixes, etc. See the alignment forum (blog) version for detailed changelogs at the bottom of each post.
011
Steve Byrnes @stevebyrnes.bsky.social · 08/08/2025
If you too would like to be falsely accused of AI ghostwriting from how effortlessly and fluently you can touch-type em dashes and other unicode glyphs… then check out my handy guide! [It’s from a decade ago, but I keep it updated.] sjbyrnes.com/unicode.html
sjbyrnes.com
Touch-Typing Unicode: How and Why
Let’s say I want to type the character μ. I look up the shortcut on the cheat sheet, and see that it’s “[compose key] * m”. So if the compose key is Right-Alt (for example), I would press and release Right-Alt, then *, then m. And μ appears!
131
Steve Byrnes @stevebyrnes.bsky.social · 06/08/2025
I went on a podcast! lironshapira.substack.com/p/the-man-wh...
lironshapira.substack.com
The Man Who Might SOLVE AI Alignment — Dr. Steven Byrnes, AGI Safety Researcher @ Astera Institute
Dr. Steven Byrnes, UC Berkeley physics PhD and Harvard physics postdoc, is an AI safety researcher at the Astera Institute, focused on solving the technical AI alignment problem.
010
Steve Byrnes @stevebyrnes.bsky.social · 05/08/2025
New blog post: “The perils of under- vs over-sculpting AGI desires”. (1/5) www.alignmentforum.org/posts/grgb2i...
120
Steve Byrnes @stevebyrnes.bsky.social · 23/07/2025
New blog post: “Behaviorist” RL Reward Functions Lead To Scheming. I argue that, if RL is used to push AI capabilities towards AGI, it will eventually lead to AI that “schemes” (feigns niceness, while looking for a chance for escape, world takeover, etc.) (1/3) www.alignmentforum.org/posts/FNJF3S...
alignmentforum.org
“Behaviorist” RL reward functions lead to scheming — AI Alignment Forum
I will argue that a large class of reward functions, which I call “behaviorist”, and which includes almost every reward function in the RL and LLM literature, are all doomed to eventually lead to AI t...
152