Sign in

Adam Gleave

@gleave.me
176 followers 61 following 39 posts

CEO & co-founder @far.ai non-profit | PhD from Berkeley | Alignment & robustness

PostsRepliesMedia
Reposted by Adam Gleave
FAR.AI @far.ai · 13/02/2026
1. Can you trust models trained directly against probes? We train an LLM against a deception probe and find four outcomes: honesty, blatant deception, obfuscated policy (fools the probe via text), or obfuscated activations (fools it via internal representations).
153
Adam Gleave @gleave.me · 03/02/2026
Excited to see the new International AI Safety Report come out! In a world of AI hype the Report cuts through to highlight where capabilities have advanced and lagged, and surveys risks in a nuanced evidence-based way. Recommended.
000
Adam Gleave @gleave.me · 13/01/2026
Look forward to presenting our STACK attack at #AAAI2026 that we've used to bypass safeguards in frontier models like GPT-5 and Opus 4
020
Adam Gleave @gleave.me · 02/07/2025
With SOTA defenses LLMs can be difficult even for experts to exploit. Yet developers often compromise on defenses to retain performance (e.g. low-latency). This paper shows how these compromises can be used to break models – and how to securely implement defenses.
120
Adam Gleave @gleave.me · 11/06/2025
So many great talks from the Singapore Alignment Workshop -- I look forward to catching up on those that I missed in person!
040
Adam Gleave @gleave.me · 04/06/2025
As I say in the video, innovation vs safety is a false dichotomy -- do check out great ideas from our speakers for how innovation can enable effective policy in the video 👇 and initial talk recordings!
000
Adam Gleave @gleave.me · 06/05/2025
AI control is one of the most exciting new research directions; excited to have the videos from ControlConf, the world's first control-specific conference. Tons of great material both intros & diving into specific areas!
020
Reposted by Adam Gleave
FAR.AI @far.ai · 30/04/2025
AI security needs more than just testing, it needs guarantees. Evan Miyazono calls for broader adoption of formal proofs, suggesting a new paradigm where AI produces code to meet human specifications.
132
Adam Gleave @gleave.me · 25/04/2025
Had a great time at the Singapore Alignment Workshop earlier this week -- fantastic start to the ICLR week! My only complaint is I missed many of the excellent talks because I was having so many interesting conversations. Looking forward to the videos to catch up!
020
Reposted by Adam Gleave
FAR.AI @far.ai · 28/03/2025
ControlConf 2025 Day 2 delivered! From robust evals to security R&D & moral patienthood, we covered the edge of AI control theory and practice. Thanks to Ryan Greenblatt, Rohin Shah, Alex Mallen, Stephen McAleer, Tomek Korbak, Steve Kelly & others for their insights.
121
Adam Gleave @gleave.me · 20/03/2025
My biggest complaint with the AI Security Forum was too much great content across the three tracks. Looking forward to catching up on the talks I missed with the videos 👇
000
Adam Gleave @gleave.me · 13/03/2025
Excited to meet others working on or interested in alignment at the Alignment Workshop Open Social before ICLR!
000
Adam Gleave @gleave.me · 11/03/2025
Excited to see people before ICLR at Alignment Workshop Singapore!
000
Adam Gleave @gleave.me · 10/03/2025
Humans sometimes cheat at exams -- might AIs do the same? Unique challenge to evaluating intelligent systems.
000
Adam Gleave @gleave.me · 03/03/2025
AI agents can start VMs, buy things, send e-mails, etc. AI control is a promising way to prevent harmful agent actions -- whether by accident, due to adversarial attack, or the systems themselves trying to subvert controls. Apply to the world's first control conference 👇
010
Adam Gleave @gleave.me · 24/02/2025
Since joining FAR.AI in June, Lindsay has delivered amazing events like the Alignment Workshop Bay Area and Paris Security Forum. Welcome to the team!
far.ai
FAR.AI
FAR.AI works to ensure AI systems are trustworthy and beneficial to society.
020
Adam Gleave @gleave.me · 17/02/2025
Evaluations are key to understanding AI capabilities and risks -- but which ones matter? Enjoyed Soroush's talk exploring these issues!
000
Adam Gleave @gleave.me · 14/02/2025
Excited to have Annie join our team, and help produce a 200-person event in her first month! We're growing across operations and technical roles -- check out opportunities 👇
000
Adam Gleave @gleave.me · 10/02/2025
I had great conversations at the AI Security Forum -- it's exciting to see people from cybersec, hardware root of trust, and AI come together to come up with creative solutions to boost AI security.
000
Adam Gleave @gleave.me · 31/01/2025
Formal verification has a lot of exciting applications, especially in the age of LLMs: e.g. can LLMs output programs with proofs of correctness? However formally verifying neural network behavior in general seems intractable -- enjoyed Zac's talk on limitations.
000
Adam Gleave @gleave.me · 28/01/2025
Many eyes on code secures critical open-source code -- we similarly need independent scrutiny of AI models to catch issues and have trust in the models. A safe harbor for evaluation as proposed by @shaynelongpre.bsky.social could enable an independent testing ecosystem.
010
Adam Gleave @gleave.me · 23/01/2025
Deceptive alignment is one of the more pernicious safety risks. I used to find it far fetched but LLMs are very good at persuasion so have the capability -- it's just a question of whether it's incentivized during training. Great to see work towards detecting deception!
010
Adam Gleave @gleave.me · 21/01/2025
Gradient routing can influence where a neural network learns a particular skill -- useful for interpretability and control.
000
Adam Gleave @gleave.me · 17/01/2025
Open-weight models are more flexible, decentralize power in AI and power research. But their flexibility also allows them to be misused: e.g. the surge in gen AI phishing. Tamper-resistant safeguards could give best of both worlds! Challenging but very important research area.
000
Reposted by Adam Gleave
FAR.AI @far.ai · 09/01/2025
"Balancing helpfulness and safety is a real challenge for AI agents in mobile environments." Kimin Lee from KAIST introduces MobileSafetyBench—a tool for testing AI safety in mobile devices.
111
Reposted by Adam Gleave
FAR.AI @far.ai · 10/01/2025
🛡️🔐 Paris AI Security Forum 2025 on Feb 9! 🇫🇷 Join 150+ engineers, researchers & policymakers from leading labs, academia & government to discuss challenges in AI system security & safety. Hear from speakers at RAND, Palo Alto Networks, Dreadnode, Pattern Labs & more. 🔗👇
121
Reposted by Adam Gleave
FAR.AI @far.ai · 13/01/2025
“We found that if you ask the LLM, surprisingly it always says that I'm 100% confident about my reasoning.” Chirag Agarwal examines the (un)reliability of chain-of-thought reasoning, highlighting issues in faithfulness, uncertainty & hallucination.
111
Adam Gleave @gleave.me · 11/01/2025
Looking forward to the AI Security forum on Feb 9th -- just before Paris AI Action Summit and two days after IASEAI. AI opens up new classes of security vulnerabilities and AI model weights may be one of the most valuable assets making AI infosec timely & fascinating!
020
Adam Gleave @gleave.me · 07/01/2025
Empirical testbeds for ways in which AI systems can misbehave are crucial to make technical progress; great to hear from Evan on his work on model organisms.
000
Adam Gleave @gleave.me · 26/12/2024
Scalable oversight is key to aligning models on tasks where AIs are better than most (or all) evaluators -- excited to see @_julianmichael_'s empirical work on this!
000
Reposted by Adam Gleave
FAR.AI @far.ai · 18/12/2024
If we want AI policy informed by science, we must build rigorous, third-party model evaluation tools. Stephen Casper advocates for model manipulation attacks to reveal hidden AI risks, especially for open-weight and closed-source models.
121
Reposted by Adam Gleave
FAR.AI @far.ai · 19/12/2024
"Are models capable of serious harm still stupid enough to fall prey to simple jailbreaks?" Adam Gleave examines if scaling alone can solve AI vulnerabilities. While adversarial training boosts resilience, it’s still orders of magnitude less efficient than attacks.
111
Reposted by Adam Gleave
FAR.AI @far.ai · 12/12/2024
China classifies AI safety as a national security issue with cybersecurity, biological security & natural disasters. Kwan Yee Ng outlined China’s policies: model registration, safety checks for gen AI, and AGI safety pilots in Beijing, Shanghai, etc. #AlignmentWorkshop
132
Reposted by Adam Gleave
FAR.AI @far.ai · 13/12/2024
Reminder: Come visit us at #NeurIPS2024 11am-2pm today! 🔬 InterpBench: Semi-Synthetic Transformers for Evaluating Mechanistic Interpretability Techniques 🎓 Rohan Gupta · Iván Arcuschin Moreno · Thomas Kwa · Adrià Garriga-Alonso
neurips.cc
NeurIPS Poster InterpBench: Semi-Synthetic Transformers for Evaluating Mechanistic Interpretability Techniques
Abstract: Mechanistic interpretability methods aim to identify the algorithm a neural network implements, but it is difficult to validate such methods when the algorithm is unknown. This work…
041
Adam Gleave @gleave.me · 12/12/2024
I often get asked how to stay in the loop on what FAR.AI is working on. We now have a a newsletter! 📷 This will be low volume (once a quarter) summarizing our key achievements and a sneak peak of future plans. Sign up 👇 to keep updated!
000
Adam Gleave @gleave.me · 10/12/2024
Experiments are critical for rigorous interpretability methods making causally valid predictions: enjoyed Atticus's talk on limitations of current methods and ways to improve them.
000
Adam Gleave @gleave.me · 09/12/2024
Excited to have four papers at NeurIPS across reward overoptimization, generalization, & inteerpretability. Won't be making it to NeurIPS myself this year but do chat to my colleagues Adrià Garriga-Alonso and Aaron Tucker if you're attending!
000
Adam Gleave @gleave.me · 05/12/2024
Optimized misalignment is one of the AI threat models with the greatest potential risk but has seen relatively little work investigating it -- loved this overview by Anca on this topic.
010
Adam Gleave @gleave.me · 04/12/2024
Knowing what AI systems can and can't do is crucial to making informed development, deployment and governance decisions. Really enjoyed Beth of @metr.org's summary of their recent evaluations work!
021