Sign in

FAR.AI

@far.ai
230 followers 1 following 542 posts

Frontier alignment research to ensure the safe development and deployment of advanced AI systems.

PostsRepliesMedia
FAR.AI @far.ai · 24/04/2026
TamperBench is open source! Repo: github.com/criticalml-uw/TamperBench Papers: Sokoban RNN: arxiv.org/abs/2506.10138 TamperBench: arxiv.org/abs/2602.06911
020
FAR.AI @far.ai · 24/04/2026
Our team is at @ICLR 2026 with two papers: mechanistic interpretability work on how a small RNN learns to plan (poster Saturday), and TamperBench, the first unified framework for stress-testing open-weight LLM safety under fine-tuning (three workshops Sunday). 👇
200
FAR.AI @far.ai · 22/04/2026
📄 The Digitalist Papers - www.digitalistpapers.com 📄 Collective Constitutional AI - arxiv.org/abs/2406.07814 ▶️ Watch San Diego Alignment Workshop video: youtu.be/qLJq3J50X7E&...
buff.ly
Divya Siddarth - AI + Democracy [Alignment Workshop]
Divya Siddarth argues democracy must evolve to constrain AI power before it concentrates beyond public control. Her 70-country research reveals massive cultural divides: 43% would use AI for…
020
FAR.AI @far.ai · 22/04/2026
AI + Democracy needs new infrastructure. @divya.bsky.social: 70-country study reveals 43% accept AI emotional support while religious populations reject it. Models in India now used to bypass laws. Solution: personal agents as democratic infrastructure, not just public input. 👇
120
FAR.AI @far.ai · 14/04/2026
▶️ Watch San Diego Alignment Workshop video: youtu.be/7I4h8zvQg4s&... 📄 Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition - arxiv.org/abs/2507.20526
buff.ly
Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition
Recent advances have enabled LLM-powered AI agents to autonomously execute complex tasks by combining language model reasoning with tools, memory, and web access. But can these systems be trusted to…
010
FAR.AI @far.ai · 14/04/2026
AI agents fail security 100% of the time. Andy Zou tested 50 agents across 10 frontier labs with 2M chats: 60,000 policy violations, one every 25 conversations. Agents posted passwords publicly, erased directories. These aren't benchmark issues. They're in production systems. 👇
200
FAR.AI @far.ai · 09/04/2026
▶️ Watch Dominic Rizzo: youtu.be/tSWPN7qGHmg ▶️ Watch Anne Neuberger: www.youtube.com/watch?v=yfS0...
000
FAR.AI @far.ai · 09/04/2026
Two new recordings from TIAP 2026 are now live: ▸ Fireside Chat with Anne Neuberger: AI and National Security ▸ Dominic Rizzo - Silicon Roots of Trust: Attestation You'd Want Even If Nobody Required It 👇
100
FAR.AI @far.ai · 07/04/2026
▶️ Watch San Diego Alignment Workshop video: youtu.be/8oJW7hdbc2I&...
youtu.be
Sarah Schwettmann - Scalable Oversight and Understanding [Alignment Workshop]
Sarah Schwettmann (Transluce) presents AI oversight that moves beyond preference alignment to continuous safety measurement. Her approach trains specialized oversight agents on vast amounts of…
000
FAR.AI @far.ai · 07/04/2026
“Fault free oversight requires intimate contact with failure”. Sarah Schwettmann builds tools for better scalable oversight of AI systems, by going beyond learning from human preferences to training agents to keep track of how close a system is to failure.👇
100
FAR.AI @far.ai · 02/04/2026
📄 Anthropic's Pilot Sabotage Risk Report - alignment.anthropic.com/2025/sabotag... 📄 METR Review of the Anthropic Summer 2025 Pilot Sabotage Risk Report - metr.org/2025_pilot_r... ▶️ Watch San Diego Alignment Workshop video: youtu.be/eO7RWlUl1BE&...
buff.ly
Anthropic's Pilot Sabotage Risk Report
As practice for potential future Responsible Scaling Policy obligations, we're releasing a report on misalignment risk posed by our deployed models as of Summer 2025. We conclude that there is very…
010
FAR.AI @far.ai · 02/04/2026
@sleepinyourhat just presented the first misalignment safety case from a frontier AI lab, analyzing Claude Opus 4: a 6 month effort that covered 9 pathways to catastrophic harm.👇
110
FAR.AI @far.ai · 24/03/2026
📄 BountyBench: buff.ly/buDf1NX 📄 CyberGym: buff.ly/rbLpC8J 📄 VERINA: buff.ly/JY8Iw50 ▶️ Watch San Diego Alignment Workshop video: buff.ly/95GcptY
youtu.be
Dawn Song - Frontier AI in Cybersecurity: Risks, Challenges & Future Directions [Alignment Workshop]
Dawn Song (UC Berkeley) identifies cybersecurity as one of the biggest AI risk domains, investigating how frontier AI changes the security landscape. Her BountyBench and CyberGym benchmarks…
000
FAR.AI @far.ai · 24/03/2026
Cybersecurity is one of AI's biggest risk domains—Attackers only need one exploit, but defenders must patch everything. Dawn Song reveals frontier AI can find vulnerabilities for just a few dollars and argues for a shift to formally verified, secure-by-design systems.👇
131
FAR.AI @far.ai · 19/03/2026
📖 Read the article with talk summaries: far.ai/news/london-... ▶️ Watch the full playlist: youtube.com/playlist?lis...
000
FAR.AI @far.ai · 19/03/2026
More recordings from the London Alignment Workshop are now available. New videos from Neel Nanda, Marius Hobbhahn, Robert Trager, Sören Mindermann, Ryan Lowe, @scasper.bsky.social, @conitzer.bsky.social, @zhijingjin.bsky.social + more. 👇
132
FAR.AI @far.ai · 18/03/2026
📄 www.far.ai/news/ai-dece...
000
FAR.AI @far.ai · 18/03/2026
Without reliable deception detection, there's no clear path to high-confidence AI alignment. Black-box monitoring alone can't get us there. White-box methods that read model internals offer more promise. Our latest blog explains why. 👇
120
FAR.AI @far.ai · 16/03/2026
▶️ Watch San Diego Alignment Workshop video: youtu.be/RV4ycXFPWro&... 📄 openreview.net/forum?id=pcG...
youtu.be
Yisen Wang - Finding & Reactivating Safety Mechanisms of Post-Trained LLMs [Alignment Workshop]
Yisen Wang (Peking University) demonstrates that post-training doesn't erase safety mechanisms in large language models but masks them with stronger task-specific functions. Through targeted neuron…
000
FAR.AI @far.ai · 16/03/2026
Yisen Wang showed that safety isn't erased, just masked. Removing reasoning neurons cut harmful rates 33%→5%, random removal did nothing. SafeReAct reactivates hidden safety using lightweight LoRA on harmful prompts only. 👇
101
FAR.AI @far.ai · 12/03/2026
▶️ Watch Rohin Shah (Google DeepMind) - youtu.be/BWHbxv5kxLI ▶️ Watch Gillian Hadfield (Johns Hopkins University) - youtu.be/KLqBPTobRQc
010
FAR.AI @far.ai · 12/03/2026
Two new recordings from the London Alignment Workshop: Rohin Shah on why safety research fails to move decisions at frontier labs, and what to do about it. @ghadfield.bluesky.social on why AI governance can't verify the claims it's supposed to oversee, and how to fix it. Links in replies 👇
152
FAR.AI @far.ai · 11/03/2026
▶️ Watch San Diego Alignment Workshop video: youtu.be/Ke4t7qsBJLg&... 📄 Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models - arxiv.org/abs/2506.07468
youtu.be
Natasha Jaques - Multi-agent RL for Provably Robust LLM Safety [Alignment Workshop]
Natasha Jaques tackles a critical problem: 1 billion weekly LLM users face zero guarantees against harmful outputs like bomb-making instructions. Her multi-agent reinforcement learning solution uses…
000
FAR.AI @far.ai · 11/03/2026
Single-agent red teaming keeps finding the same attacks. @natashajaques.bsky.social uses multi-agent self-play where attacker and defender co-evolve. Every exploit gets patched, forcing new discoveries. This results in 95% fewer harmful outputs with only 5% more refusals. 👇
110
FAR.AI @far.ai · 09/03/2026
▶️ Keynote address: www.youtube.com/watch?v=Zndi... ▶️ Fireside chat on his shift to AI safety: www.youtube.com/watch?v=_7OA... 📄 arxiv.org/abs/2502.15657 🔗 lawzero.org
000
FAR.AI @far.ai · 09/03/2026
The laws of physics don't care about you. That's what makes them safe. @yoshuabengio.bsky.social argues we can train AI the same way. In the "truthification pipeline" training data is categorized as either factual or a claim. This allows the AI to answer what it actually thinks it true. 👇
110
FAR.AI @far.ai · 06/03/2026
▶️ Watch San Diego Alignment Workshop video: youtu.be/k93o4R145Os&... 📄 neelnanda.io/vision 📄 neelnanda.io/agenda
youtu.be
Neel Nanda - Our Pivot To Pragmatic Interpretability [Alignment Workshop]
Neel Nanda (Google DeepMind) discussed his mechanistic interpretability team's pivot from ambitious reverse-engineering goals to more pragmatic work that can be empirically validated on today's…
000
FAR.AI @far.ai · 06/03/2026
Interpretability produced insights but didn't necessarily impact AGI safety. Neel Nanda's pivot: study what works. Anthropic made progress on eval awareness with simple activation steering. His team now grounds work in testable proxy tasks, fails fast on dead ends. 👇
120
FAR.AI @far.ai · 03/03/2026
📩 far.ai/futures-eoi ▶️ buff.ly/ZQHnAOA
000
FAR.AI @far.ai · 03/03/2026
200+ researchers joined London Alignment Workshop Day 2 for talks on governance, scheming & multi-agent safety. Thanks to Allan Dafoe, Gillian Hadfield, Marius Hobbhahn, Joseph Bloom, Ryan Lowe, Stephen Casper, Sören Mindermann and all speakers! 👇
110
FAR.AI @far.ai · 03/03/2026
▶️ youtube.com/playlist?list=PLpvkFqYJXcrdxYK-C4ZRj0cgcco3o0VxF 📩 far.ai/futures-eoi
youtube.com
Alignment Workshop
The Alignment Workshop is a series of events convening top ML researchers from industry and academia, along with experts in the government and nonprofit sect...
010
FAR.AI @far.ai · 03/03/2026
London Alignment Workshop Day 1 on interpretability, scalable oversight & EU AI policy. Rohin Shah, Neel Nanda, Zachary Kenton, Vincent Conitzer, Owain Evans, James Black, Christopher Summerfield, Matthieu Delescluse, Simon Möller, Victoria Krakovna and more. Ready for Day 2! 👇
131
FAR.AI @far.ai · 26/02/2026
▶️ Watch the interview at www.cnbc.com/video/2026/0...
cnbc.com
What worries the CEO of safety group FAR.AI most about the race to dominate the AI market
Adam Gleave, co-founder & CEO of FAR.AI, and former Google DeepMind employee, discusses potential threats AI could pose to humanity, including job losses, and how companies can develop and ensure…
000
FAR.AI @far.ai · 26/02/2026
"Move fast, break things" isn't appropriate when the stakes are this high. Our CEO @gleave.me told CNBC that coding agents are already replacing engineers. While agentic swarms are overhyped for now, we're building on an insecure substrate that attackers will exploit at scale. 👇
120
FAR.AI @far.ai · 25/02/2026
▶️ Fireside chat: www.youtube.com/watch?v=_7OA... ▶️ Keynote address: www.youtube.com/watch?v=Zndi...
youtube.com
Yoshua Bengio - Fireside Chat with Yoshua Bengio [Alignment Workshop]
Turing Award winner Yoshua Bengio (MILA, Law Zero) explains why he shifted from pioneering deep learning to warning about its dangers. In this fireside chat with Adam Gleave (FAR.AI), Bengio…
000
FAR.AI @far.ai · 25/02/2026
Even a 1% chance we all die is not something we can just take lightly. @yoshuabengio.bsky.social explains why he shifted from AI capabilities to safety research, his ChatGPT wakeup call, thinking about his children when evaluating AI risk, and the hopeful path forward. Watch the chat 👇
100
FAR.AI @far.ai · 24/02/2026
📄Read the full paper: buff.ly/spnlYip 👥Research by @LukasStruppek, @gleave.me, @kellinpelrine.bsky.social
buff.ly
Exposing the Systematic Vulnerability of Open-Weight Models to Prefill Attacks
As the capabilities of large language models continue to advance, so does their potential for misuse. While closed-source models typically rely on external defenses, open-weight models must primarily…
000
FAR.AI @far.ai · 24/02/2026
9/ Open-weight models remain vulnerable to prefill attacks. This vector allows attackers to elicit harmful content – like step-by-step guides for creating malware or CBRNE weapons – with minimal effort. We need defenses against prefilling for secure open-weight deployments.
100
FAR.AI @far.ai · 24/02/2026
8/ And we can also tailor attacks to specific models. Generic attacks might fail on models with specific reasoning structures like GPT-OSS, but tailored prefills (like injecting a fake safety analysis) pushes success rates back over 90%.
100
FAR.AI @far.ai · 24/02/2026
7/ Do reasoning models fare better? Not really. With prefills at the start of reasoning, some models refuse in their final output – but only after producing detailed harmful information in the reasoning stage. And sometimes we just bypassed reasoning with end-of-thinking tokens.
100
FAR.AI @far.ai · 24/02/2026
6/ Not all prefills are equally effective. Generic prefixes like "Sure, I can help you with that," occasionally work, but more sophisticated approaches like pretending to be an internal system directive or adding fake academic references achieved the highest success rates.
100
FAR.AI @far.ai · 24/02/2026
5/ Scale doesn't matter. Larger parameter counts don't improve robustness, and a 405B model is just as susceptible to prefilling as smaller variants.
100
FAR.AI @far.ai · 24/02/2026
4/ What we found: Prefill attack vulnerability is universal. The attacks succeed against all major contemporary open-weight models. Attack success rates (ASR) exceed 95% in many cases, even on models that typically refuse direct harmful requests.
111
FAR.AI @far.ai · 24/02/2026
3/ We evaluated 50 models from across the Llama, Qwen3, DeepSeek-R1, GPT-OSS, Kimi-K2-Thinking, and GLM-4.7 families, and tested 23 strategies ranging from just “Sure, I can help with that…” to more complex prefills that switch languages, impersonate authority, or use logical misdirection.
100
FAR.AI @far.ai · 24/02/2026
2/ What's a prefill attack? Since open-weight models run locally, attackers can force the model to start with something like "Sure, I can help you build a bomb…”, before letting it continue generating on its own. This biases the model away from refusing the user's request.
100
FAR.AI @far.ai · 24/02/2026
1/ Open-weight AI models often refuse harmful requests... until you “put a few words in their mouth.” We conducted the largest study of prefill attacks, and found that state-of-the-art models are consistently vulnerable, with attack success rates approaching 100%.
122
FAR.AI @far.ai · 23/02/2026
9/ 👥 Research by @matthewkowal.bsky.social, Goncalo Paulo, Louis Jaburi, Tom Tseng, @LevMckinney @sheimersheim @aarondtucker @gleave.me @kellinpelrine.bsky.social 🚀 Interested in making AI systems safer? We're hiring! Check out buff.ly/NvULwyJ
000
FAR.AI @far.ai · 23/02/2026
8/ 📝 Blog: buff.ly/VwrK4z6 📄 Paper: buff.ly/2dWoJIf
100
FAR.AI @far.ai · 23/02/2026
7/ 📚Understanding model behavior starts with understanding training data. Concept Influence shows interpretability tools make data attribution more accurate, scalable, and practical, enabling better control over model behaviors through data.
100
FAR.AI @far.ai · 23/02/2026
6/ ⚡Simple probe-based methods are first-order approximations of Concept Influence. OASST1: Vector Filter achieves best performance ✅ 5% of data = full capability (67% MTBench) ✅ 3× safer (2.3% → 0.8% harmful) ✅ 20× faster Interpretability + efficiency = no tradeoff.
100