Sign in

Fazl Barez

@fbarez.bsky.social
261 followers 144 following 78 posts

Let's build AI's we can trust! fazlbarez.com

PostsRepliesMedia
Reposted by Fazl Barez
University of Oxford @ox.ac.uk · 22/09/2026
AI is advancing fast. But reliably measuring what it can do remains difficult. @oxmartinschool.bsky.social's Prof Maike Osborne and Dr Fazl Barez explore why better ways to evaluate and monitor AI matter for the decisions ahead ⬇️ bit.ly/4yf6vTI
 Graphic titled “Flying Blind into the AI Age”. A glowing digital brain is surrounded by streams of data and computer-like nodes. Text highlights expert commentary from Professor Maike Osborne and Dr Fazl Barez on the uncertainty around AI, the difficulty of measuring its capabilities, and the need to make high-stakes decisions about AI soon.
073
Fazl Barez @fbarez.bsky.social · 21/09/2026
@maosbot.bsky.social and I shared our thoughts on whats going on with AI and what we are doing to help!
000
Fazl Barez @fbarez.bsky.social · 18/09/2026
As AI gets more autonomy, we need stronger evidence that we can keep it under control. I was glad to join @dwdderwetterdienst.bsky.social news to discuss unexpected AI behaviour and what it means for safety. We still have agency over AIs and we should only build AGI if we can align them!
000
Fazl Barez @fbarez.bsky.social · 08/09/2026
I’m excited to give an invited talk at eXCV workshop @eccv.bsky.social ! For the best part of history we relied on understanding to arrive at answers. Yet, AI's can produce answers we cannot understand fully. In my talk i'll discuss what understanding may look like in the age of AGI!
I will explore the changing relationship because for the best part of history we relied on understanding to arrive at answers. Yet, Increasingly, AI's can produce answers we cannot understand (at least fully).
020
Fazl Barez @fbarez.bsky.social · 03/09/2026
I was glad to join Sky News to discuss some of the safety questions around humanoid robots. The technology is exciting, but as AI systems begin acting in the physical world, it becomes even more important to test them carefully and make sure people remain in control.
110
Fazl Barez @fbarez.bsky.social · 04/07/2026
I'll be at #ICML2026 next week—7 main conference papers, 2 orals + an invited talk 🇰🇷 Also hiring 2 RAs (interp + continual learning) at Oxford: what happens inside models that keep learning after deployment? Come chat in Seoul, or apply 👇 tsglab.github.io/vacancies/ Please share with folks!
tsglab.github.io
Vacancies | TSG Lab – Technical Safety & Governance Lab
010
Reposted by Fazl Barez
Oxford Internet Institute @oii.ox.ac.uk · 02/07/2026
A glimpse into another successful Oxford Connected Life Summit, focused on what it means to live in an increasingly connected world. This year's theme was "New Intelligence, Old Questions," and featured notable speakers and organisations. Huge thanks to the student committee for making this happen!
011
Fazl Barez @fbarez.bsky.social · 30/06/2026
Heading to #ICML2026 in Seoul next week with the TSG Lab and Martian 🇰🇷 10 papers: 7 in the main conf, 3 WS: interpretability, AI evaluation, and governance, two oral spotlights Giving an invited talk at the EIML WS Grateful to the students, collaborators, mentors and Claude who made it happen!
000
Fazl Barez @fbarez.bsky.social · 25/06/2026
Really grateful to have 7 papers accepted at @icmlconf.bsky.social onf 2026, including 2 spotlights! Massive thanks to all my collaborators—I’ve been lucky to work with such brilliant people #ICML2026
100
Fazl Barez @fbarez.bsky.social · 25/06/2026
Excited to be debating at the Oxford Union this evening Motion: This House Believes that AI is the Great Equalizer Is it? Or isn't it? I'm speaking for the proposition--which might surprise those who know my work. That's rather the point! We'll find out which way the House votes
001
Reposted by Fazl Barez
Stella Biderman @stellaathena.bsky.social · 10/06/2026
In film, "we'll fix it in post" is what you say when something went wrong on set and you don't want to redo it. AI research has made it our entire methodology: train the model, then patch whatever comes out. Our new ICML oral argues this can't be the basis of a science of AI. 🧵
310823
Reposted by Fazl Barez
Oxford Martin School @oxmartinschool.bsky.social · 09/06/2026
How can we ensure AI-powered robots remain safe when operating in the real world? 🤖 A recent article co-authored by @aigioxfordmartin.bsky.social researcher Fazl Barez, explores safety challenges and the importance of being context-aware 🔗Read the research in full: www.science.org/doi/10.1126/...
science.org
Beyond alignment: Why robotic foundation models need context-aware safety
Because AI-enabled robots can be tricked into taking unsafe actions, they require layered, context-aware safety guardrails.
031
Fazl Barez @fbarez.bsky.social · 07/12/2025
Incredibly excited to announce $1 Million prize pool to solve the world’s most important scientific problem in Interpretability. The goal is to turns hard interpretability questions into tools for human empowerment, oversight and governance.
241
Fazl Barez @fbarez.bsky.social · 06/10/2025
🚨New AI Safety Course @aims_oxford ! I’m thrilled to launch a new called AI Safety & Alignment (AISAA) course on the foundations & frontier research of making advanced AI systems safe and aligned at @UniofOxford what to expect 👇 robots.ox.ac.uk/~fazl/aisaa/
161
Reposted by Fazl Barez
Toby Ord @tobyord.bsky.social · 25/09/2025
Evaluating the Infinite 🧵 My latest paper tries to solve a longstanding problem afflicting fields such as decision theory, economics, and ethics — the problem of infinities. Let me explain a bit about what causes the problem and how my solution avoids it. 1/N arxiv.org/abs/2509.19389
arxiv.org
Evaluating the Infinite
I present a novel mathematical technique for dealing with the infinities arising from divergent sums and integrals. It assigns them fine-grained infinite values from the set of hyperreal numbers in a ...
2125
Fazl Barez @fbarez.bsky.social · 19/09/2025
🚀 Excited to have 2 papers accepted at #NeurIP2025! 🎉 congrats to my amazing co-authors! More details (and more bragging) soon! and maybe even more news on sep 25 👀 See you all in… Mexico? San Diego? Copenhagen? Who knows! 🌍✈️
010
Reposted by Fazl Barez
Jakob Mökander @jakobmokander.bsky.social · 04/09/2025
🚨 NEW PAPER 🚨: Embodied AI (incl. AI-powered drones, self-driving cars and robots) is here, but policies are lagging. We analyzed the EAI risks and found significant gaps in governance arxiv.org/pdf/2509.00117 Co-authors Jared Perlo @fbarez.bsky.social Alex Robey & @floridi.bsky.social 1\4
133
Reposted by Fazl Barez
Martin Tutek @mtutek.bsky.social · 21/08/2025
Other works have highlighted that CoTs ≠ explainability alphaxiv.org/abs/2025.02 (@fbarez.bsky.social), and that intermediate (CoT) tokens ≠ reasoning traces arxiv.org/abs/2504.09762 (@rao2z.bsky.social). Here, FUR offers a fine-grained test if LMs latently used information from CoTs for answers!
alphaxiv.org
Chain-of-Thought Is Not Explainability | alphaXiv
View 3 comments: There should be a balance of both subjective and observable methodologies. Adhering to just one is a fools errand.
161
Reposted by Fazl Barez
Jeroen ‘Jeremy’ Fransen @jeroenjeremy.bsky.social · 02/07/2025
It is so easy to confuse chain of thought and explainability and in fact in a lot of the media it is presented as if with current LLMs we are allowed to view their actual thought processes. It is not that!
062
Fazl Barez @fbarez.bsky.social · 01/07/2025
Excited to share our paper: "Chain-of-Thought Is Not Explainability"! We unpack a critical misconception in AI: models explaining their steps (CoT) aren't necessarily revealing their true reasoning. Spoiler: the transparency can be an illusion. (1/9) 🧵
28531
Fazl Barez @fbarez.bsky.social · 27/06/2025
Technology = power. AI is reshaping power — fast. Today’s AI doesn’t just assist decisions; it makes them. Governments use it for surveillance, prediction, and control — often with no oversight. Technical safeguards aren’t enough on their own — but they’re essential for AI to serve society.
140
Reposted by Fazl Barez
David Duvenaud @davidduvenaud.bsky.social · 18/06/2025
And Anna Yelizarov, @fbarez.bsky.social, @scasper.bsky.social, Beatrice Erkers, among others. We'll draw from political theory, cooperative AI, economics, mechanism design, history, and hierarchical agency.
131
Reposted by Fazl Barez
Yoav Gur Arieh @yoav.ml · 29/05/2025
This is a step toward targeted, interpretable, and robust knowledge removal — at the parameter level. Joint work with Clara Suslik, Yihuai Hong, and @fbarez.bsky.social, advised by @megamor2.bsky.social 🔗 Paper: arxiv.org/abs/2505.22586 🔗 Code: github.com/yoavgur/PISCES
011
Fazl Barez @fbarez.bsky.social · 20/05/2025
Come work with me at Oxford this summer! Paid research opportunity to: White-box LLMs & model security Safe RL & reward hacking Interpretability & governance tools Remote or Oxford. Apply by 30 May 23:59 UTC. DM with questions.
120
Fazl Barez @fbarez.bsky.social · 15/05/2025
Come work with me at Oxford! We’re hiring a Postdoc in Causal Systems Modelling to: - Build causal & white-box models that make frontier AI safer and more transparent - Turn technical insights into safety cases, policy briefs, and governance tools ] DM if you have any questions.
144
Fazl Barez @fbarez.bsky.social · 08/04/2025
First-time Area Chair seeking advice! What helped you most when evaluating papers beyond just averaging scores? After suffering through unhelpful reviews as an author, I want to do right by papers in my track.
100
Reposted by Fazl Barez
Mor Geva @megamor2.bsky.social · 31/03/2025
🎉 Our Actionable Interpretability workshop has been accepted to #ICML2025! 🎉 > Follow @actinterp.bsky.social > Website actionable-interpretability.github.io @talhaklay.bsky.social @anja.re @mariusmosbach.bsky.social @sarah-nlp.bsky.social @iftenney.bsky.social Paper submission deadline: May 9th!
34116
Reposted by Fazl Barez
Technical AI Governance @ ICML 2025 @taig-icml.bsky.social · 01/04/2025
Organizers: Ben Bucknall, @lisasoder.bsky.social, @ankareuel.bsky.social @fbarez.bsky.social, @carlosmougan.bsky.social Weiwei Pan, Siddharth Swaroop, @ankareuel.bsky.social , Robert Trager @maosbot.bsky.social
052
Fazl Barez @fbarez.bsky.social · 01/04/2025
Technical AI Governance (TAIG) at #ICML2025 this July in Vancouver! Credit to Ben and Lisa for all the work! We have a new centre at Oxford working on technical AI governance with Robert Trager and @maosbot.bsky.social many other great minds. We are hiring - please reach out! Quote
061
Reposted by Fazl Barez
Naomi Saphra @nsaphra.bsky.social · 27/03/2025
Life update: I'm starting as faculty at Boston University @bucds.bsky.social in 2026! BU has SCHEMES for LM interpretability & analysis, I couldn't be more pumped to join a burgeoning supergroup w/ @najoung.bsky.social @amuuueller.bsky.social. Looking for my first students, so apply and reach out!
CDS building which looks like a jenga tower
3524213
Reposted by Fazl Barez
Itay Itzhak @ COLM 🍁 @itay-itzhak.bsky.social · 17/03/2025
New paper alert! Curious how small prompt tweaks impact LLM accuracy but don’t want to run endless inferences? We got you. Meet DOVE - a dataset built to uncover these sensitivities. Use DOVE for your analysis or contribute samples -we're growing and welcome you aboard!
041
Reposted by Fazl Barez
wdmacaskill.bsky.social @wdmacaskill.bsky.social · 17/03/2025
What happens once AI can design better AI, which can itself design better AI? Will we get an "intelligence explosion" where AI capabilities increase very rapidly? Tom Davidson, Rose Hadshar and I have a new paper out with analysis of these dynamics.
151
Reposted by Fazl Barez
Jakob Foerster @jfoerst.bsky.social · 12/03/2025
My group @FLAIR_Ox is recruiting a postdoc and looking for someone who can get started by the end of April. Deadline to apply is in one week (!), 19th of March at noon, so please help spread the word: my.corehr.com/pls/uoxrecru...
my.corehr.com
Job Details
01913
Reposted by Fazl Barez
Tal Haklay @talhaklay.bsky.social · 06/03/2025
1/13 LLM circuits tell us where the computation happens inside the model—but the computation varies by token position, a key detail often ignored! We propose a method to automatically find position-aware circuits, improving faithfulness while keeping circuits compact. 🧵👇
1268
Fazl Barez @fbarez.bsky.social · 04/03/2025
🔍 Excited to share our paper: "Same Question, Different Words: A Latent Adversarial Framework for Prompt Robustness"!
231
Fazl Barez @fbarez.bsky.social · 01/03/2025
New paper alert! 🚨 Important question: Do SAEs generalise? We explore the answerability detection in LLMs by comparing SAE features vs. linear residual stream probes. Answer: probes outperform SAE features in-domain, out-of-domain generalization varies sharply between features and datasets. 🧵
1102
Reposted by Fazl Barez
Adi Simhi @adisimhi.bsky.social · 19/02/2025
🚨New arXiv preprint!🚨 LLMs can hallucinate - but did you know they can do so with high certainty even when they know the correct answer? 🤯 We find those hallucinations in our latest work with @itay-itzhak.bsky.social, @fbarez.bsky.social, @gabistanovsky.bsky.social and Yonatan Belinkov
32110
Reposted by Fazl Barez
Oxford Martin AI Governance Initiative @aigioxfordmartin.bsky.social · 18/02/2025
We are excited to welcome Fazl Barez @fbarez.bsky.social, who joins us as a senior postdoctoral research fellow. He will be leading research initiatives in AI safety and interpretability. @oxmartinschool.bsky.social Find out more: www.oxfordmartin.ox.ac.uk/people/fazl-...
163
Reposted by Fazl Barez
Yoshua Bengio @yoshuabengio.bsky.social · 11/01/2025
Very interesting paper about unlearning for AI Safety, a subject that deserves more attention. ⬇️
0506
Fazl Barez @fbarez.bsky.social · 10/01/2025
🚨 New Paper Alert: Open Problem in Machine Unlearning for AI Safety 🚨 Can AI truly "forget"? While unlearning promises data removal, controlling emergent capabilities is a inherent challenge. Here's why it matters: 👇 Paper: arxiv.org/pdf/2501.04952 1/8
1256
Reposted by Fazl Barez
rylanschaeffer.bsky.social @rylanschaeffer.bsky.social · 13/12/2024
What happens when "If at first you don't succeed, try again?" meets modern ML/AI insights about scaling up? You jailbreak every model on the market😱😱😱 Fire work led by @jplhughes.bsky.social Sara Price @aengusl.bsky.social Mrinank Sharma Ethan Perez arxiv.org/abs/2412.03556
032
Reposted by Fazl Barez
jplhughes.bsky.social @jplhughes.bsky.social · 06/12/2024
🚨🛡️ Jailbreak Defense in a Narrow Domain 🛡️🚨 Jailbreaking is easy. Defending is hard. Might defending against a single, narrow behavior be easier? Even in this focused setting, all defenses fail 😱 arxiv.org/abs/2412.02159 Appearing at @AdvMLFrontiers (Oral) & @solarneurips #NeurIPS2024
arxiv.org
Jailbreak Defense in a Narrow Domain: Limitations of Existing...
Defending large language models against jailbreaks so that they never engage in a broadly-defined set of forbidden behaviors is an open problem. In this paper, we investigate the difficulty of...
241
Fazl Barez @fbarez.bsky.social · 04/12/2024
Today is a good day for AI Safety! We are launching the AI Luminate AI Safety Benchmark @MLCommons @PeterMattson100 @tangenticAI The first step towards global standard benchmark for AI PRODUCT safety!
010
Reposted by Fazl Barez
Maike Osborne @maosbot.bsky.social · 03/12/2024
Our new Chancellor wants us admission tutors to "vet" Chinese applicants. How the heck are we supposed to do that? We're underpaid and overworked academics, not a national intelligence service www.politico.eu/article/oxfo...
politico.eu
Oxford University’s China dilemma
Politicians are among those vying to run the elite institution — and they’re squaring off over Beijing’s influence on Britain.
5445
Fazl Barez @fbarez.bsky.social · 02/12/2024
ML /AI Safety researchers how do you find time to read papers, work on your projects and be so active on b sky, x etc?
000
Fazl Barez @fbarez.bsky.social · 22/11/2024
🚨 Join us at NeurIPS 2024! 🚨 🧠 PrivacyML: Meaningful Privacy-Preserving ML & Evaluations 🌐 privacy.github.io Tackling the questions in AI Privacy & Safety: How do we protect training data privacy? What does unlearning mean for AI safety? Can cryptography make AI safer?
privacy.github.io
120