Sign in

Fazl Barez

@fbarez.bsky.social
264 followers 144 following 79 posts

Let's build AI's we can trust! fazlbarez.com

PostsRepliesMedia
Fazl Barez @fbarez.bsky.social · 18/09/2026
As AI gets more autonomy, we need stronger evidence that we can keep it under control. I was glad to join @dwdderwetterdienst.bsky.social news to discuss unexpected AI behaviour and what it means for safety. We still have agency over AIs and we should only build AGI if we can align them!
000
Fazl Barez @fbarez.bsky.social · 08/09/2026
I’m excited to give an invited talk at eXCV workshop @eccv.bsky.social ! For the best part of history we relied on understanding to arrive at answers. Yet, AI's can produce answers we cannot understand fully. In my talk i'll discuss what understanding may look like in the age of AGI!
I will explore the changing relationship because for the best part of history we relied on understanding to arrive at answers. Yet, Increasingly, AI's can produce answers we cannot understand (at least fully).
020
Fazl Barez @fbarez.bsky.social · 03/09/2026
I was glad to join Sky News to discuss some of the safety questions around humanoid robots. The technology is exciting, but as AI systems begin acting in the physical world, it becomes even more important to test them carefully and make sure people remain in control.
110
Fazl Barez @fbarez.bsky.social · 30/06/2026
Heading to #ICML2026 in Seoul next week with the TSG Lab and Martian 🇰🇷 10 papers: 7 in the main conf, 3 WS: interpretability, AI evaluation, and governance, two oral spotlights Giving an invited talk at the EIML WS Grateful to the students, collaborators, mentors and Claude who made it happen!
000
Fazl Barez @fbarez.bsky.social · 25/06/2026
Excited to be debating at the Oxford Union this evening Motion: This House Believes that AI is the Great Equalizer Is it? Or isn't it? I'm speaking for the proposition--which might surprise those who know my work. That's rather the point! We'll find out which way the House votes
001
Fazl Barez @fbarez.bsky.social · 07/12/2025
The theme of the prize is Interpretability for Code Generation. Why? Because code offers ground truth. Unlike natural language, code is formal and allows us to measure faithfulness and track progress
100
Fazl Barez @fbarez.bsky.social · 07/12/2025
We are taking a skeptical lens to the current state of the art. Research needs to answer the hard questions holding the field back - aka Actionable Are our methods Scalable? → Are they Complete? → And most critically—Are they actually useful for fix things?
100
Fazl Barez @fbarez.bsky.social · 07/12/2025
Incredibly excited to announce $1 Million prize pool to solve the world’s most important scientific problem in Interpretability. The goal is to turns hard interpretability questions into tools for human empowerment, oversight and governance.
241
Fazl Barez @fbarez.bsky.social · 06/10/2025
We’ll cover: 1️⃣ The alignment problem -- foundations & present-day challenges 2️⃣ Frontier alignment methods & evaluation (RLHF, Constitutional AI, etc.) 3️⃣ Interpretability & monitoring incl. hands-on mech interp labs 4️⃣ Sociotechnical aspects of alignment, governance, risks, Economics of AI and policy
120
Fazl Barez @fbarez.bsky.social · 06/10/2025
🚨New AI Safety Course @aims_oxford ! I’m thrilled to launch a new called AI Safety & Alignment (AISAA) course on the foundations & frontier research of making advanced AI systems safe and aligned at @UniofOxford what to expect 👇 robots.ox.ac.uk/~fazl/aisaa/
161
Fazl Barez @fbarez.bsky.social · 01/07/2025
@alasdair-p.bsky.social‬, @adelbibi.bsky.social ‬, Robert Trager, Damiano Fornasiere, @john-yan.bsky.social ‬, @yanai.bsky.social@yoshuabengio.bsky.social
220
Fazl Barez @fbarez.bsky.social · 01/07/2025
Bottom line: CoT can be useful but should never be mistaken for genuine interpretability. Ensuring trustworthy explanations requires rigorous validation and deeper insight into model internals, which is especially critical as AI scales up in high-stakes domains. (9/9) 📖✨
140
Fazl Barez @fbarez.bsky.social · 01/07/2025
Language models can be prompted or trained to verbalize reasoning steps in their Chain of Thought (CoT). Despite prior work showing such reasoning can be unfaithful, we find that around 25% of recent CoT-centric papers still mistakenly claim CoT as an interpretability technique. (2/9)
170
Fazl Barez @fbarez.bsky.social · 01/07/2025
Excited to share our paper: "Chain-of-Thought Is Not Explainability"! We unpack a critical misconception in AI: models explaining their steps (CoT) aren't necessarily revealing their true reasoning. Spoiler: the transparency can be an illusion. (1/9) 🧵
28531
Fazl Barez @fbarez.bsky.social · 27/06/2025
Technology = power. AI is reshaping power — fast. Today’s AI doesn’t just assist decisions; it makes them. Governments use it for surveillance, prediction, and control — often with no oversight. Technical safeguards aren’t enough on their own — but they’re essential for AI to serve society.
140
Fazl Barez @fbarez.bsky.social · 04/03/2025
🔍 Excited to share our paper: "Same Question, Different Words: A Latent Adversarial Framework for Prompt Robustness"!
231
Fazl Barez @fbarez.bsky.social · 01/03/2025
💡 Feature Similarity: Cosine similarities reveal that linear probes learn highly similar features, while high-performing SAE features are often dissimilar. While individual SAE features are also dissimilar to the probes, combinations of SAE features are increasingly similar. 6/
100
Fazl Barez @fbarez.bsky.social · 01/03/2025
🔍 Digging Deeper: We analyzed combinations of SAE features. While increasing the number of features improves the in-domain performance (SQUAD), it doesn’t consistently improve generalization performance when compared to the best single SAE features. 5/
100
Fazl Barez @fbarez.bsky.social · 01/03/2025
📊 Out-of-Distribution Insights: Performance varies across domains: • We find SAE features with strong generalization to the Equation, Celebrity, or IDK dataset, but no good generalization to BoolQ • The probe trained on BoolQ fails to generalize, although decent SAE features exist 4/
100
Fazl Barez @fbarez.bsky.social · 01/03/2025
🛠️ Our Approach: We compare two probing methods at different layers (notably Layer 31): • Linear probes trained on the residual stream, hitting ~90% accuracy in-domain • Individual SAE features, reaching ~80% accuracy 3/
100
Fazl Barez @fbarez.bsky.social · 01/03/2025
🔍 The Challenge: LLMs sometimes confidently answer questions—even ones they can’t truly answer. We use three existing datasets (SQuAD, IDK, BoolQ), and two newly created synthetic datasets (Equations, and Celebrity Names) to quantify this behavior. 2/
110
Fazl Barez @fbarez.bsky.social · 01/03/2025
New paper alert! 🚨 Important question: Do SAEs generalise? We explore the answerability detection in LLMs by comparing SAE features vs. linear residual stream probes. Answer: probes outperform SAE features in-domain, out-of-domain generalization varies sharply between features and datasets. 🧵
1102
Fazl Barez @fbarez.bsky.social · 10/01/2025
⚠️ Current unlearning tests fall short. Models can pass tests yet retain harmful, emergent capabilities. These risks often surface in subtle, unexpected ways during extended interactions. How can we evaluate major concerns e.g. the dual-use capabilities can emerge from beneficial data? 4/8
110
Fazl Barez @fbarez.bsky.social · 10/01/2025
💡 Key AI Safety Applications for Unlearning: Managing safety-critical knowledge Mitigating jailbreaks Correcting value alignment Ensuring privacy/legal compliance But: Removing facts is easy; removing dangerous capabilities is HARD. Capabilities emerge from complex knowledge interactions. 3/8
140
Fazl Barez @fbarez.bsky.social · 10/01/2025
There's a crucial distinction between removing specific knowledge vs controlling capabilities. While you can make an AI "forget" certain facts, preventing it from reconstructing capabilities is much harder since they emerge from combining benign knowledge. Real world example: 2a/8
110
Fazl Barez @fbarez.bsky.social · 10/01/2025
🚨 New Paper Alert: Open Problem in Machine Unlearning for AI Safety 🚨 Can AI truly "forget"? While unlearning promises data removal, controlling emergent capabilities is a inherent challenge. Here's why it matters: 👇 Paper: arxiv.org/pdf/2501.04952 1/8
1256
Fazl Barez @fbarez.bsky.social · 04/12/2024
Today is a good day for AI Safety! We are launching the AI Luminate AI Safety Benchmark @MLCommons @PeterMattson100 @tangenticAI The first step towards global standard benchmark for AI PRODUCT safety!
010