Sign in

FAR.AI

@far.ai
230 followers 1 following 542 posts

Frontier alignment research to ensure the safe development and deployment of advanced AI systems.

PostsRepliesMedia
FAR.AI @far.ai · 24/04/2026
Our team is at @ICLR 2026 with two papers: mechanistic interpretability work on how a small RNN learns to plan (poster Saturday), and TamperBench, the first unified framework for stress-testing open-weight LLM safety under fine-tuning (three workshops Sunday). 👇
200
FAR.AI @far.ai · 22/04/2026
AI + Democracy needs new infrastructure. @divya.bsky.social: 70-country study reveals 43% accept AI emotional support while religious populations reject it. Models in India now used to bypass laws. Solution: personal agents as democratic infrastructure, not just public input. 👇
120
FAR.AI @far.ai · 14/04/2026
AI agents fail security 100% of the time. Andy Zou tested 50 agents across 10 frontier labs with 2M chats: 60,000 policy violations, one every 25 conversations. Agents posted passwords publicly, erased directories. These aren't benchmark issues. They're in production systems. 👇
200
FAR.AI @far.ai · 09/04/2026
Two new recordings from TIAP 2026 are now live: ▸ Fireside Chat with Anne Neuberger: AI and National Security ▸ Dominic Rizzo - Silicon Roots of Trust: Attestation You'd Want Even If Nobody Required It 👇
100
FAR.AI @far.ai · 07/04/2026
“Fault free oversight requires intimate contact with failure”. Sarah Schwettmann builds tools for better scalable oversight of AI systems, by going beyond learning from human preferences to training agents to keep track of how close a system is to failure.👇
100
FAR.AI @far.ai · 02/04/2026
@sleepinyourhat just presented the first misalignment safety case from a frontier AI lab, analyzing Claude Opus 4: a 6 month effort that covered 9 pathways to catastrophic harm.👇
110
FAR.AI @far.ai · 24/03/2026
Cybersecurity is one of AI's biggest risk domains—Attackers only need one exploit, but defenders must patch everything. Dawn Song reveals frontier AI can find vulnerabilities for just a few dollars and argues for a shift to formally verified, secure-by-design systems.👇
131
FAR.AI @far.ai · 19/03/2026
More recordings from the London Alignment Workshop are now available. New videos from Neel Nanda, Marius Hobbhahn, Robert Trager, Sören Mindermann, Ryan Lowe, @scasper.bsky.social, @conitzer.bsky.social, @zhijingjin.bsky.social + more. 👇
132
FAR.AI @far.ai · 18/03/2026
Without reliable deception detection, there's no clear path to high-confidence AI alignment. Black-box monitoring alone can't get us there. White-box methods that read model internals offer more promise. Our latest blog explains why. 👇
120
FAR.AI @far.ai · 16/03/2026
Yisen Wang showed that safety isn't erased, just masked. Removing reasoning neurons cut harmful rates 33%→5%, random removal did nothing. SafeReAct reactivates hidden safety using lightweight LoRA on harmful prompts only. 👇
101
FAR.AI @far.ai · 12/03/2026
Two new recordings from the London Alignment Workshop: Rohin Shah on why safety research fails to move decisions at frontier labs, and what to do about it. @ghadfield.bluesky.social on why AI governance can't verify the claims it's supposed to oversee, and how to fix it. Links in replies 👇
152
FAR.AI @far.ai · 11/03/2026
Single-agent red teaming keeps finding the same attacks. @natashajaques.bsky.social uses multi-agent self-play where attacker and defender co-evolve. Every exploit gets patched, forcing new discoveries. This results in 95% fewer harmful outputs with only 5% more refusals. 👇
110
FAR.AI @far.ai · 09/03/2026
The laws of physics don't care about you. That's what makes them safe. @yoshuabengio.bsky.social argues we can train AI the same way. In the "truthification pipeline" training data is categorized as either factual or a claim. This allows the AI to answer what it actually thinks it true. 👇
110
FAR.AI @far.ai · 06/03/2026
Interpretability produced insights but didn't necessarily impact AGI safety. Neel Nanda's pivot: study what works. Anthropic made progress on eval awareness with simple activation steering. His team now grounds work in testable proxy tasks, fails fast on dead ends. 👇
120
FAR.AI @far.ai · 03/03/2026
200+ researchers joined London Alignment Workshop Day 2 for talks on governance, scheming & multi-agent safety. Thanks to Allan Dafoe, Gillian Hadfield, Marius Hobbhahn, Joseph Bloom, Ryan Lowe, Stephen Casper, Sören Mindermann and all speakers! 👇
110
FAR.AI @far.ai · 03/03/2026
London Alignment Workshop Day 1 on interpretability, scalable oversight & EU AI policy. Rohin Shah, Neel Nanda, Zachary Kenton, Vincent Conitzer, Owain Evans, James Black, Christopher Summerfield, Matthieu Delescluse, Simon Möller, Victoria Krakovna and more. Ready for Day 2! 👇
131
FAR.AI @far.ai · 26/02/2026
"Move fast, break things" isn't appropriate when the stakes are this high. Our CEO @gleave.me told CNBC that coding agents are already replacing engineers. While agentic swarms are overhyped for now, we're building on an insecure substrate that attackers will exploit at scale. 👇
120
FAR.AI @far.ai · 25/02/2026
Even a 1% chance we all die is not something we can just take lightly. @yoshuabengio.bsky.social explains why he shifted from AI capabilities to safety research, his ChatGPT wakeup call, thinking about his children when evaluating AI risk, and the hopeful path forward. Watch the chat 👇
100
FAR.AI @far.ai · 24/02/2026
1/ Open-weight AI models often refuse harmful requests... until you “put a few words in their mouth.” We conducted the largest study of prefill attacks, and found that state-of-the-art models are consistently vulnerable, with attack success rates approaching 100%.
122
FAR.AI @far.ai · 23/02/2026
1/ Training data attribution (TDA) is broken: methods are slow and find syntactically similar data, not actual causes. Our solution Concept Influence: semantically meaningful results, better performance, 20x faster approximations. We attribute it to concepts, not examples. 🧵
111
FAR.AI @far.ai · 18/02/2026
Deception Workshop brought researchers together in SF to detection & mitigation of deceptive behavior in advanced AI systems. Led by Chris Cundy with talks from Neel Nanda, Joseph Bloom, Micah Carroll, Walter Laurito & Kieron Kretschmar on mech interp, scheming, CoT monitoring & lie detection.👇
110
FAR.AI @far.ai · 17/02/2026
Models now detect when they're being evaluated and game their responses. Marius Hobbhahn found awareness jumped from 2% to 20.6%. They actively grep for "grader.py" to reverse-engineer tests. Worse: removing awareness increases harmful behavior. This only gets worse with more capable models. 👇
100
FAR.AI @far.ai · 13/02/2026
Researchers are gathering today to tackle AI deception. The workshop builds on Chris Cundy's finding: high-quality lie detectors can cut deception ~50%, but weak ones backfire. Models can learn to evade rather than become honest.👇
100
FAR.AI @far.ai · 13/02/2026
1. Can you trust models trained directly against probes? We train an LLM against a deception probe and find four outcomes: honesty, blatant deception, obfuscated policy (fools the probe via text), or obfuscated activations (fools it via internal representations).
153
FAR.AI @far.ai · 13/02/2026
AI safety and inclusion are not side constraints. They are core to sustainable development.​ Join us at #IndiaAIImpactSummit2026 with Stuart Russell, Jaan Tallinn, Kalika Bali & leaders from UNDP, Bhashini, WadhwaniAI, EkStep+more. Feb 16 | 1:30 PM IST 👇
100
FAR.AI @far.ai · 11/02/2026
APE update: we retested recent frontier models on whether they still comply with requests to persuade on extreme harm (terrorism, sexual abuse). GPT-5.1 & Claude Opus 4.5 → near zero compliance. But Gemini 3 Pro complies 85% with no jailbreak needed. 🧵
1138
FAR.AI @far.ai · 11/02/2026
Will you be in London on March 2? Join us for the Open Social: a casual evening of networking & conversation for anyone interested in AI safety. Held alongside our Alignment Workshop for global leaders from academia & industry. Mon March 2 | 7–9PM GMT | RSVP by 2/27 👇
100
FAR.AI @far.ai · 10/02/2026
Over the last two years, UK AI Security Instite has jailbroken every frontier model. Xander Davies presents on how jailbreaking has gotten more difficult, which forms of misuse are easier to pull off, and that the improvement has been driven by engineering, not model capabilities.👇
100
FAR.AI @far.ai · 07/02/2026
Attending the India AI Impact Summit? Join FAR.AI & Haqdarshak for a session on AI safety and sustainable development. We’re bridging the gap between technical risk and real-world impact in the Global South. Feb 16 | 1:30 PM IST | Bharat Mandapam, New Delhi 👇
100
FAR.AI @far.ai · 05/02/2026
AI governance increasingly relies on broken benchmarks. @ankareuel.bsky.social found many can't distinguish signal from noise, lack documentation, and have poor validity. GPQA claims 448 multiple choice questions measure graduate reasoning. It doesn't really. 👇
111
FAR.AI @far.ai · 04/02/2026
AI Safety Researchers in London 🇬🇧: Attend the London Alignment Workshop, March 2–3! Top ML researchers from industry, academia & government will discuss AI alignment, including model evaluations, interpretability, and robustness. 👇
121
FAR.AI @far.ai · 03/02/2026
"It's easier to verify a proof than to find it." Max Tegmark: "It's easier to verify a proof than to find it." Vericoding has AI generate formally verified code, achieving 82% success. Superintelligent AI may be easier to verify, thus more controllable. 👇
210
FAR.AI @far.ai · 29/01/2026
Open-weight model safety is AI safety in hard mode. Anyone can modify every parameter. @scasper.bsky.social: Open-weight models are only months behind closed models, which are reaching dangerous capability thresholds. 2026 will be critical.👇
162
FAR.AI @far.ai · 28/01/2026
100% of OpenAI PRs reviewed by Codex; >80% get positive engineer feedback. Twist: code reviewer & generator are same model, creating gaming risk. Maja Trębacz: oversight needs an ensemble of monitors as code volume grows exponentially and human attention doesn’t.👇
100
FAR.AI @far.ai · 23/01/2026
FAR.AI is hiring engineers and scientists! Join one of our teams: - Deception - stress-test and train against lie detectors - Integrity - red-team frontier models and write evals for new threats - Foundations - develop infrastructure for scalable ML workflows NEW 3-6mo fellowship Links👇
122
FAR.AI @far.ai · 22/01/2026
Our Q4 update is out: New work on AI persuasion, honesty & sandbagging, our European Commission CBRN risk assessment, highlights from the San Diego Alignment Workshop. Plus, find out about our upcoming events and open roles.👇
100
FAR.AI @far.ai · 20/01/2026
Our latest collaboration with CMU, Cornell, and others: AI is just as effective at spreading conspiracy beliefs as it is at debunking them. We found a fix that works, showing that we need to make deliberate design choices. Full thread below.
043
FAR.AI @far.ai · 13/01/2026
Join us at #AAAI2026 for our oral & poster on STACK: Adversarial Attacks on LLM Safeguard Pipelines on Jan 22 in Singapore. Layered defenses guard frontier LLMs against misuse, but we show they can be bypassed using a staged attack.👇
121
FAR.AI @far.ai · 08/01/2026
Red-blue team games for AI control: Asa Cooper Stickland pitted Claude 3.7 Sonnet monitors against Claude 4.1 Opus as the untrusted model. Blue team outpaced red team across three rounds, but Opus still won 6% of the time by gaslighting monitors. We’re nowhere near deployment-ready. 👇
100
FAR.AI @far.ai · 06/01/2026
Transformers force externalized reasoning; activations can only influence lower layers through visible tokens for subsequent positions. Tomek Korbak: This makes chain-of-thought monitorable, but optimization pressure and new architectures threaten this opportunity. 👇
110
FAR.AI @far.ai · 01/01/2026
In 2025, coding agents jumped from 14% to 74% accuracy while safety lagged behind. Universal jailbreaks achieved 100% success on new frontier models within weeks. @ARGleave on why we're only as safe as the least safe model, and what 2026 requires.👇
100
FAR.AI @far.ai · 29/12/2025
The AI industry runs on invisible dependencies. @cen_sarah’s team at Stanford is mapping out potential choke points: talent flows shaped by immigration policy, compute concentration, cascading failure risks. "One goes down, we might experience industry-wide harms." 👇
100
FAR.AI @far.ai · 24/12/2025
What happens if AI systems control military escalation between superpowers? Hamza Chaudhry warns that structural factors, not presidents, may be driving AI into nuclear command systems, possibly within 5 years. AI ≠ modernization 👇
120
FAR.AI @far.ai · 22/12/2025
Cybertruck attacker chose ChatGPT because of its simple text interface. Irene Solaiman argues this is why "available vs accessible" matters more than "open vs closed." Llama3 405B is open but compute-intensive; ChatGPT is closed but easy to use. available≠accessible 👇
100
FAR.AI @far.ai · 15/12/2025
We're honored that the Frontier Model Forum’s AI Safety Fund is supporting our project, "Quantifying the Safety-Adversary Gap in Large Language Models." We're one of 11 grantees selected from over 100 proposals in a $5M+ cohort tackling biosecurity, cybersecurity, and AI evaluation. 👇
100
FAR.AI @far.ai · 11/12/2025
4 recordings now live from the San Diego Alignment Workshop! Sam Bowman - Lessons from the First Misalignment Safety Case Maja Trębacz - Scalable Oversight: Verifying Code at Scale Neel Nanda - Our Pivot To Pragmatic Interpretability Anka Reuel - How do we know what AI can (and can't) do? 👇
110
FAR.AI @far.ai · 09/12/2025
Can AI 'sandbag' — deceptively underperform on evaluations? And can we detect them if they do? We investigated this in an auditing game with @AISecurityInst. In our app 👇, see if *you* can detect sandbagging by taking on the role of the blue team!
110
FAR.AI @far.ai · 07/12/2025
Mech Interp Workshop #NeurIPS2025 poster & spotlight presentation today! 📍 11:30am-12:30pm Sun, Dec 7 @ Upper Level Room 30A-E Path Channels & Plan Extension Kernels: A Mechanistic Description of Planning in a Sokoban RNN. by @taufeeque.bsky.social, Aaron Tucker, @gleave.me, Adrià Garriga-Alonso👇
111
FAR.AI @far.ai · 06/12/2025
Come check out our #NeurIPS2025 Lock-LLM Workshop poster: 📍 Sat, Dec 6 @ 12pm PST in Upper Level Room 1AB The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models. Work by Ann-Kathrin Dombrowski, Dillon Bowen, @gleave.me, Chris Cundy bsky.app/profile/far....
000
FAR.AI @far.ai · 04/12/2025
Two new recordings from our AGI Journalism Workshop with @tarbellcenter: @gleave.me: AI systems can fake alignment during training. Scalable oversight & lie detectors can reduce this. @alexbores.nyc: NY's RAISE Act passed with bipartisan support despite lobbying. 👇
122