Sign in

Kyle O’Brien

@kyletokens.bsky.social
54 followers 105 following 12 posts

AGI Alignment & Control Pretraining @ Geodesic Research | kyobrien.io

PostsRepliesMedia
Reposted by Kyle O’Brien
Geodesic Research @geodesicresearch.bsky.social · 28/05/2026
Thanks to a generous philanthropic grant (pending final logistics) from Coefficient Giving, 𝘎𝘦𝘰𝘥𝘦𝘴𝘪𝘤 𝘪𝘴 𝘩𝘪𝘳𝘪𝘯𝘨 𝘔𝘦𝘮𝘣𝘦𝘳𝘴 𝘰𝘧 𝘛𝘦𝘤𝘩𝘯𝘪𝘤𝘢𝘭 𝘚𝘵𝘢𝘧𝘧. Come build the base of alignment with us 🤖 Applications now open: airtable.com/appuugUGFPJE...
142
Kyle O’Brien @kyletokens.bsky.social · 19/05/2026
Geodesic mission is to develop the science of providing robustly aligned initializations for RL, where alignment priors persist through the remainder of training. Do considering applying if you want to help make sharing the world with superintelligence go well for humanity.
000
Kyle O’Brien @kyletokens.bsky.social · 02/02/2026
Very excited to see pretraining safety efforts! We’re only now beginning to understand how promising pretraining safety and alignment interventions are. Much in the way that curating the base model is important for capabilities like reasoning, so too might it be important for safety.
030
Kyle O’Brien @kyletokens.bsky.social · 16/01/2026
I've joined Geodesic Research to build the open-science field of AI safety pretraining research. Our first paper is wild. TL;DR — LLMs pretrained on data about misaligned AIs themselves become less aligned. Luckily, pretraining LLMs with data about good AIs helps them become more aligned.
alignmentpretraining.ai
Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment
LLMs trained on data about misaligned AIs themselves become less aligned. Luckily, pretraining LLMs with synthetic data about good AIs helps them become more aligned. These alignment priors persist th...
040
Reposted by Kyle O’Brien
Alex Turner @turntrout.bsky.social · 21/12/2025
The first pretraining results are in, and it looks like models indeed have self-fulfilling misalignment properties. Great work by Tice et al! alignmentpretraining.ai
alignmentpretraining.ai
Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment
LLMs trained on data about misaligned AIs themselves become less aligned. Luckily, pretraining LLMs with synthetic data about good AIs helps them become more aligned. These alignment priors persist…
061
Kyle O’Brien @kyletokens.bsky.social · 14/12/2025
You know you're AGI-pilled when your Spotify Wrapped looks like this.
010
Kyle O’Brien @kyletokens.bsky.social · 31/10/2025
Applications to apply for the ERA:AI Fellowship close November 3rd! Participating in this Summer's fellowship was my gateway into pursuing AGI safety research full-time. I will be a research manager for the upcoming Winter fellowships. Feel free to DM me with questions. :) erafellowship.org
erafellowship.org
ERA Fellowship
ERA is a talent programme supporting early-career researchers and entrepreneurs to understand and mitigate risks from frontier AI, based at Cambridge, UK.
000
Reposted by Kyle O’Brien
Cas (Stephen Casper) @scasper.bsky.social · 04/09/2025
📌📌📌 I'm excited to be on the faculty job market this fall. I just updated my website with my CV. stephencasper.com
stephencasper.com
Stephen Casper
Visit the post for more.
0184
Reposted by Kyle O’Brien
Sharon Goldman @sharongoldman.bsky.social · 15/08/2025
Thanks to @stellaathena.bsky.social for chatting with me about Deep Ignorance: the new paper/project from Eleuther AI and the UK AISI. Bottom line: Worried AI could teach people to build bioweapons? Don’t teach it how fortune.com/2025/08/14/w...
fortune.com
AI safety tip: if you don’t want it giving bioweapon instructions, maybe don’t put them in the training data, say researchers
New research shows that scrubbing risky material from AI training data can build safeguards that are harder to bypass — and one author calls out tech giants for keeping such work under wraps.
0112
Kyle O’Brien @kyletokens.bsky.social · 12/08/2025
This articles covers our work for a general audience. :)
040
Kyle O’Brien @kyletokens.bsky.social · 12/08/2025
Big and True :)
030
Kyle O’Brien @kyletokens.bsky.social · 10/08/2025
I like that OpenAI published this. They were able to fine-tune away GPT-oss's refusal, decreasing refusal rates to ~0%. These results aren't surprising. Acknowledging that existing safeguards don't generalize to open models is the first step in developing solutions. arxiv.org/abs/2508.031...
arxiv.org
Estimating Worst-Case Frontier Risks of Open-Weight LLMs
In this paper, we study the worst-case frontier risks of releasing gpt-oss. We introduce malicious fine-tuning (MFT), where we attempt to elicit maximum capabilities by fine-tuning gpt-oss to be as ca...
010
Kyle O’Brien @kyletokens.bsky.social · 02/08/2025
I've learned a lot over the past two years of getting into research, mostly from mistakes. I’ve made many mistakes. Such is science. Good research is often at the adjacent possible. I've written up much of what I've learned now that I'm beginning to mentor others. open.substack.com/pub/kyletoke...
open.substack.com
Don’t "Think", Just Think
Lessons From Breaking Into AI Research
000
Kyle O’Brien @kyletokens.bsky.social · 20/06/2025
I led an effort at Microsoft last Fall that studied whether SAE steering was an effective way to improve jailbreak robustness. Our paper on SAE steering has been accepted to the ICML Actionable Interpretability Workshop! Venue: actionable-interpretability.github.io Paper: arxiv.org/abs/2411.11296
arxiv.org
Steering Language Model Refusal with Sparse Autoencoders
Responsible deployment of language models requires mechanisms for refusing unsafe prompts while preserving model performance. While most approaches modify model weights through additional training, we...
020
Kyle O’Brien @kyletokens.bsky.social · 07/06/2025
I'll be in England this summer as an AI Safety Research Fellow with ERA! erafellowship.org/fellowship I will be studying data filtering and tamper-resistant unlearning for open-weight AI safety so that the community can continue to benefit from open models as capabilities improve.
erafellowship.org
Fellowship — ERA Fellowship
150