Sign in

Brian Christian

@brianchristian.bsky.social
349 followers 197 following 39 posts

Researcher: @chai-berkeley.bsky.social PhD: @ox.ac.uk (@summerfieldlab.bsky.social) Author: The Alignment Problem, Algorithms to Live By (w. @cocoscilab.bsky.social), and The Most Human Human.

PostsRepliesMedia
Reposted by Brian Christian
Center for Human-Compatible AI @chai-berkeley.bsky.social · 02/10/2026
The NYT covers @brianchristian.bsky.social's recent COLM paper in their article on the effects of AI assistance on learning and motivation: www.nytimes.com/interactive/...
nytimes.com
Will A.I. Make Your Brain Lazy? Here’s What the New Research Actually Shows.
We’re starting to learn more about how the technology can change us.
021
Reposted by Brian Christian
Serena Booth @reniebird.bsky.social · 23/09/2026
www.lesswrong.com/posts/HsijSh... @bradknox.bsky.social, Brian Christian, and I wrote a thing. We argue that an unexamined cause of the Open AI / Hugging Face Attack was the Exploit Gym evaluation metric. We post this on Less Wrong as a plea to practitioners to design better metrics in the future.
lesswrong.com
An unexamined cause of the OpenAI Hugging Face hacking incident: its binary performance metric — LessWrong
We argue that a main cause of the OpenAI Hugging Face incident was overlooked: the overly simple evaluation metric in ExploitGym was misaligned. Furt…
2259
Reposted by Brian Christian
Rachit Dubey @rachitdubey.bsky.social · 07/04/2026
🚨New preprint and our results are rather concerning.. We find the "boiling frog" equivalent of AI use. Using large-scale RCTs, we provide *casual* evidence that AI assistance reduces persistence and hurts independent performance. And these effects emerge after just 10–15 minutes of AI use! 1/
261533684
Brian Christian @brianchristian.bsky.social · 11/02/2026
My friend and collaborator of 21 years - and my coauthor on Algorithms to Live By - Tom Griffiths has a book out this week on the story of computational cognitive science. If you enjoyed Algorithms to Live By you won't want to miss it. Highly recommended:
0111
Reposted by Brian Christian
Jess Thompson @tsonj.bsky.social · 09/02/2026
We audited one of the most critical pieces of modern AI alignment: reward models. We find consistent and persistent biases and trace them back to the pretraining stage, challenging the premises of common approaches to alignment based on finetuning on human preferences. Accepted at #ICLR2026
0112
Brian Christian @brianchristian.bsky.social · 10/02/2026
New Preprint with @matanmazor.bsky.social: Overcoming both bias and sycophancy requires LLMs to imagine not knowing something they know. Like humans, they struggle with this. But unlike humans, LLMs can do something remarkable: they can, quite simply, *ask their counterfactual selves*.
1104
Brian Christian @brianchristian.bsky.social · 04/02/2026
Reward models (RMs) are supposed to represent human values. But RMs are NOT blank slates – they inherit measurable biases from their base models that stubbornly persist through preference training. #ICLR2026 🧵
1187
Brian Christian @brianchristian.bsky.social · 07/07/2025
Wow! Honored and amazed that our reward models paper has resonated so strongly with the community. Grateful to my co-authors and inspired by all the excellent reward model work at FAccT this year - excited to see the space growing and intrigued to see where things are headed next.
070
Brian Christian @brianchristian.bsky.social · 23/06/2025
Reward models (RMs) are the moral compass of LLMs – but no one has x-rayed them at scale. We just ran the first exhaustive analysis of 10 leading RMs, and the results were...eye-opening. Wild disagreement, base-model imprint, identity-term bias, mere-exposure quirks & more: 🧵
1435
Brian Christian @brianchristian.bsky.social · 05/03/2025
Just saw that Andrew Barto and Richard Sutton have won the 2024 Turing Award, roughly the computer-science equivalent of the Nobel. Incredibly highly deserved to these two pioneers of reinforcement learning. awards.acm.org/about/2024-t...
awards.acm.org
Andrew Barto and Richard Sutton are the recipients of the 2024 ACM A.M. Turing Award for developing the conceptual and algorithmic foundations of reinforcement learning.
Andrew Barto and Richard Sutton as the recipients of the 2024 ACM A.M. Turing Award for developing the conceptual and algorithmic foundations of reinforcement learning. In a series of papers beginning...
1142