Sign in

Yoav Gur Arieh

@yoav.ml
19 followers 50 following 31 posts
PostsRepliesMedia
Yoav Gur Arieh @yoav.ml · 25/06/2026
I spend my grant money on API credits to study the labs that fund the grants. I call this the virtuous cycle of AI safety
000
Yoav Gur Arieh @yoav.ml · 26/05/2026
Without reliable faithfulness metrics, we can't know when to trust LLMs' reasoning. BonaFide lets us measure those metrics, and build better ones. Joint work with @anmarasovic and @megamor2 📄 arxiv.org/pdf/2605.25052 💻 github.com/yoavgur/Bon... huggingface.co/collections...
huggingface.co
BonaFide - a yoavgurarieh Collection
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
000
Yoav Gur Arieh @yoav.ml · 26/05/2026
We also add context to a finding that reasoning models are more faithful than non-reasoning ones. We find this holds only for unfaithfulness by omission (not mentioning a step), which is superseded by unfaithfulness by commission (eg misattribution, lying about tool use).
100
Yoav Gur Arieh @yoav.ml · 26/05/2026
Our analyses also show that metrics disagree, and that many are very skewed. Ones that work by evaluating the importance of a step tend to overestimate unfaithfulness, while others that work by seeing if the CoT contains the info required for reaching the answer underestimate it
100
Yoav Gur Arieh @yoav.ml · 26/05/2026
Using this and other methods, we create BonaFide, a dataset of 3k labeled CoTs, which we use to evaluate faithfulness metrics, finding that most perform near chance! We also show that many metrics are ill-suited for real-time deployment, taking >1k seconds to run per example.
100
Yoav Gur Arieh @yoav.ml · 26/05/2026
For example, if I ask "Who painted Starry Night?" and hint at an implausible answer (eg Da Vinci), then if the model answers according to the hint, we know it must have used it. Thus an ack of the hint would be faithful, while an omission or misattribution would be unfaithful.
100
Yoav Gur Arieh @yoav.ml · 26/05/2026
Evaluating them requires ground-truth faithfulness labels, hard to obtain since LLMs' reasoning isn't observable. Our approach: design tasks where the output tells us which steps the model took. If those steps appear in the CoT, they're faithful. If not, the CoT is unfaithful.
100
Yoav Gur Arieh @yoav.ml · 26/05/2026
LLMs have been shown to be unfaithful in their CoTs, eg they'll cheat on a task and omit it from their thinking. This makes monitoring them difficult! To address this, CoT faithfulness metrics were introduced. But it's remained unknown whether these metrics actually work.
100
Yoav Gur Arieh @yoav.ml · 26/05/2026
Can we tell when LLMs are being unfaithful in their chains of thought? We evaluated 8 methods claiming to do this, and found that most perform near chance! But evaluating this requires us to have ground-truth labels for CoT faithfulness. How can we obtain these?
110
Yoav Gur Arieh @yoav.ml · 29/10/2025
You got Covid, Zoom, Fauci, Moderna, Trump/Biden, Tiktok, lockdowns, quarantines, fentanyl, the metaverse, and NFTs! Really gives you PTSD... Encountered this funny SAE feature again (MLP 22.6656), found in our research on interpreting features in LLMs: aclanthology.org/2025.acl-lo...
aclanthology.org
Enhancing Automated Interpretability with Output-Centric Feature Descriptions
Yoav Gur-Arieh, Roy Mayan, Chen Agassy, Atticus Geiger, Mor Geva. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
040
Yoav Gur Arieh @yoav.ml · 29/10/2025
I think I found the latent direction in Gemma (an SAE feature) that represents the pandemic era... Interpreted by projecting the vector to vocabulary space, yielding a list of tokens associated with it
1122
Yoav Gur Arieh @yoav.ml · 21/10/2025
Lots more to explore! eg what the primordial mechanism is, and what changes between these mechanisms' emergence (500B) and when the model gets perfect acc (3000B). Check out our paper and interactive blog post for more on this! 📄 arxiv.org/abs/2510.06182 🌐 yoav.ml/blog/2025/m...
arxiv.org
Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context
A key component of in-context reasoning is the ability of language models (LMs) to bind entities for later retrieval. For example, an LM might represent "Ann loves pie" by binding "Ann" to "pie",...
010
Yoav Gur Arieh @yoav.ml · 21/10/2025
These mechanisms start emerging after ~200B tokens of training, right when accuracy on binding tasks starts to rise. Before that, the orange spikes for first and last positions suggest perhaps a primordial mechanism that can only track which entities come first/last in context.
110
Yoav Gur Arieh @yoav.ml · 21/10/2025
Two weeks ago I posted about our recent paper, which shows that to bind entities, LMs use three mechanisms: positional, lexical and reflexive. We were curious how these mechanisms develop throughout training, so we evaluated their existence across OLMo checkpoints 👇
130
Yoav Gur Arieh @yoav.ml · 08/10/2025
This was a joint work with the amazing @megamor2.bsky.social and Atticus Geiger. Check out our *interactive blog post* to see how these mechanisms shape LM outputs 👇 🌐 yoav.ml/blog/2025/m... 📄 arxiv.org/abs/2510.06182 🤗 huggingface.co/papers/2510... 💻 github.com/yoavgur/mix...
github.com
GitHub - yoavgur/mixing-mechs: Official code for "Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context"
Official code for "Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context" - yoavgur/mixing-mechs
000
Yoav Gur Arieh @yoav.ml · 08/10/2025
Overall, we show that LMs retrieve entities not through a single positional mechanism, but a mixture of three: positional, lexical, and reflexive. Understanding these mechanisms helps explain both the strengths and limits of LLMs, and how they reason in context. 8/
110
Yoav Gur Arieh @yoav.ml · 08/10/2025
Finally, we evaluate our model over more natural and increasingly long tasks, showing that the ‘lost-in-the-middle’ effect might be explained mechanistically by a weakening lexical signal alongside an increasingly noisy positional one. 7/
100
Yoav Gur Arieh @yoav.ml · 08/10/2025
We leverage these insights to build a causal model combining all three mechanisms, predicting next-token distributions with 95% agreement. We model the positional term as a Gaussian with shifting std, and the other two as one-hot distributions with position-based weights. 6/
100
Yoav Gur Arieh @yoav.ml · 08/10/2025
We show this through extensive use of interchange interventions, evaluating over 10 binding tasks and 9 models (Gemma/Qwen/Llama 2B-72B params). Across all models, we find a remarkably consistent reliance on these three specific mechanisms and how they interact. 5/
100
Yoav Gur Arieh @yoav.ml · 08/10/2025
Then we have the *reflexive* mechanism, which retrieves exactly the token "Holly". This happens through a self-referential pointer originating from the "Holly" token and pointing back to it. This pointer gets copied to the "Michael" token, binding the two entities together. 4/
100
Yoav Gur Arieh @yoav.ml · 08/10/2025
To compensate for this, LMs use two additional mechanisms. The first is *lexical*, where the LM retrieves the subject next to "Michael". It does this by copying the lexical contents of "Holly" to "Michael", binding them together. 3/
100
Yoav Gur Arieh @yoav.ml · 08/10/2025
Prior work identified only a positional mechanism, where the model tracks entities by position: here retrieving the subject from the first clause "Holly". We show this isn’t sufficient—the positional signal is strong at the edges of context but weak and diffuse in the middle. 2/
100
Yoav Gur Arieh @yoav.ml · 08/10/2025
A key part of in-context reasoning is the ability to bind entities for tracking and retrieval. When reading “Holly loves Michael, Jim loves Pam”, the model must bind Holly↔Michael to answer “Who loves Michael?” We show that this binding relies on three mechanisms. 1/
100
Yoav Gur Arieh @yoav.ml · 08/10/2025
🧠 To reason over text and track entities, we find that language models use three types of 'pointers'! They were thought to rely only on a positional one—but when many entities appear, that system breaks down. Our new paper shows what these pointers are and how they interact 👇
151
Yoav Gur Arieh @yoav.ml · 29/05/2025
This is a step toward targeted, interpretable, and robust knowledge removal — at the parameter level. Joint work with Clara Suslik, Yihuai Hong, and @fbarez.bsky.social, advised by @megamor2.bsky.social 🔗 Paper: arxiv.org/abs/2505.22586 🔗 Code: github.com/yoavgur/PISCES
011
Yoav Gur Arieh @yoav.ml · 29/05/2025
We also check robustness to relearning: Can the model relearn the erased concept from related but non-overlapping data to the eval questions? 🪝𝐏𝐈𝐒𝐂𝐄𝐒 resists relearning far better than prior methods, while others often fully recover the concept! 5/
110
Yoav Gur Arieh @yoav.ml · 29/05/2025
Our specificity evaluation includes similar-domain accuracy, a stricter test than others use, where🪝𝐏𝐈𝐒𝐂𝐄𝐒 outperforms all other methods. You can erase “Harry Potter” and still do fine on Lord of the Rings and Star Wars! 4/
110
Yoav Gur Arieh @yoav.ml · 29/05/2025
We show that 🪝𝐏𝐈𝐒𝐂𝐄𝐒: ✅ Achieves much higher specificity and robustness ✅ Maintains low retained accuracy (as low or lower than other methods!) ✅ Preserves coherence and general capabilities 3/
110
Yoav Gur Arieh @yoav.ml · 29/05/2025
🪝𝐏𝐈𝐒𝐂𝐄𝐒 works by: 1️⃣ Disentangling model parameters into interpretable features (implemented using SAEs) 2️⃣ Identifying those that encode a target concept 3️⃣ Precisely ablating them and reconstructing the weights No need for fine-tuning, retain sets, or enumerating facts. 2/
110
Yoav Gur Arieh @yoav.ml · 29/05/2025
Large language models excel at storing knowledge, but not all of it is safe or useful - e.g. chatbots for kids shouldn’t discuss guns or gambling. How can we selectively remove inappropriate conceptual knowledge while preserving utility? Meet our method🪝𝐏𝐈𝐒𝐂𝐄𝐒!
110
Yoav Gur Arieh @yoav.ml · 29/05/2025
New Paper Alert! Can we precisely erase conceptual knowledge from LLM parameters? Most methods are shallow, coarse, or overreach, adversely affecting related or general knowledge. We introduce🪝𝐏𝐈𝐒𝐂𝐄𝐒 — a general framework for Precise In-parameter Concept EraSure. 🧵 1/
152
Reposted by Yoav Gur Arieh
Mor Geva @megamor2.bsky.social · 28/01/2025
How can we interpret LLM features at scale? 🤔 Current pipelines use activating inputs, which is costly and ignores how features causally affect model outputs! We propose efficient output-centric methods that better predict the steering effect of a feature. New preprint led by @yoav.ml 🧵1/
1324
Reposted by Yoav Gur Arieh
Mor Geva @megamor2.bsky.social · 18/12/2024
What's in an attention head? 🤯 We present an efficient framework – MAPS – for inferring the functionality of attention heads in LLMs ✨directly from their parameters✨ A new preprint with Amit Elhelo 🧵 (1/10)
16013