Sign in

Yoav Gur Arieh

@yoav.ml
19 followers 50 following 31 posts
PostsRepliesMedia
Yoav Gur Arieh @yoav.ml · 25/06/2026
I spend my grant money on API credits to study the labs that fund the grants. I call this the virtuous cycle of AI safety
000
Yoav Gur Arieh @yoav.ml · 26/05/2026
We also add context to a finding that reasoning models are more faithful than non-reasoning ones. We find this holds only for unfaithfulness by omission (not mentioning a step), which is superseded by unfaithfulness by commission (eg misattribution, lying about tool use).
100
Yoav Gur Arieh @yoav.ml · 26/05/2026
Our analyses also show that metrics disagree, and that many are very skewed. Ones that work by evaluating the importance of a step tend to overestimate unfaithfulness, while others that work by seeing if the CoT contains the info required for reaching the answer underestimate it
100
Yoav Gur Arieh @yoav.ml · 26/05/2026
Using this and other methods, we create BonaFide, a dataset of 3k labeled CoTs, which we use to evaluate faithfulness metrics, finding that most perform near chance! We also show that many metrics are ill-suited for real-time deployment, taking >1k seconds to run per example.
100
Yoav Gur Arieh @yoav.ml · 26/05/2026
For example, if I ask "Who painted Starry Night?" and hint at an implausible answer (eg Da Vinci), then if the model answers according to the hint, we know it must have used it. Thus an ack of the hint would be faithful, while an omission or misattribution would be unfaithful.
100
Yoav Gur Arieh @yoav.ml · 26/05/2026
Evaluating them requires ground-truth faithfulness labels, hard to obtain since LLMs' reasoning isn't observable. Our approach: design tasks where the output tells us which steps the model took. If those steps appear in the CoT, they're faithful. If not, the CoT is unfaithful.
100
Yoav Gur Arieh @yoav.ml · 26/05/2026
Can we tell when LLMs are being unfaithful in their chains of thought? We evaluated 8 methods claiming to do this, and found that most perform near chance! But evaluating this requires us to have ground-truth labels for CoT faithfulness. How can we obtain these?
110
Yoav Gur Arieh @yoav.ml · 29/10/2025
I think I found the latent direction in Gemma (an SAE feature) that represents the pandemic era... Interpreted by projecting the vector to vocabulary space, yielding a list of tokens associated with it
1122
Yoav Gur Arieh @yoav.ml · 21/10/2025
These mechanisms start emerging after ~200B tokens of training, right when accuracy on binding tasks starts to rise. Before that, the orange spikes for first and last positions suggest perhaps a primordial mechanism that can only track which entities come first/last in context.
110
Yoav Gur Arieh @yoav.ml · 21/10/2025
Two weeks ago I posted about our recent paper, which shows that to bind entities, LMs use three mechanisms: positional, lexical and reflexive. We were curious how these mechanisms develop throughout training, so we evaluated their existence across OLMo checkpoints 👇
130
Yoav Gur Arieh @yoav.ml · 08/10/2025
Finally, we evaluate our model over more natural and increasingly long tasks, showing that the ‘lost-in-the-middle’ effect might be explained mechanistically by a weakening lexical signal alongside an increasingly noisy positional one. 7/
100
Yoav Gur Arieh @yoav.ml · 08/10/2025
We leverage these insights to build a causal model combining all three mechanisms, predicting next-token distributions with 95% agreement. We model the positional term as a Gaussian with shifting std, and the other two as one-hot distributions with position-based weights. 6/
100
Yoav Gur Arieh @yoav.ml · 08/10/2025
We show this through extensive use of interchange interventions, evaluating over 10 binding tasks and 9 models (Gemma/Qwen/Llama 2B-72B params). Across all models, we find a remarkably consistent reliance on these three specific mechanisms and how they interact. 5/
100
Yoav Gur Arieh @yoav.ml · 08/10/2025
To compensate for this, LMs use two additional mechanisms. The first is *lexical*, where the LM retrieves the subject next to "Michael". It does this by copying the lexical contents of "Holly" to "Michael", binding them together. 3/
100
Yoav Gur Arieh @yoav.ml · 08/10/2025
Prior work identified only a positional mechanism, where the model tracks entities by position: here retrieving the subject from the first clause "Holly". We show this isn’t sufficient—the positional signal is strong at the edges of context but weak and diffuse in the middle. 2/
100
Yoav Gur Arieh @yoav.ml · 08/10/2025
A key part of in-context reasoning is the ability to bind entities for tracking and retrieval. When reading “Holly loves Michael, Jim loves Pam”, the model must bind Holly↔Michael to answer “Who loves Michael?” We show that this binding relies on three mechanisms. 1/
100
Yoav Gur Arieh @yoav.ml · 08/10/2025
🧠 To reason over text and track entities, we find that language models use three types of 'pointers'! They were thought to rely only on a positional one—but when many entities appear, that system breaks down. Our new paper shows what these pointers are and how they interact 👇
151
Yoav Gur Arieh @yoav.ml · 29/05/2025
This is a step toward targeted, interpretable, and robust knowledge removal — at the parameter level. Joint work with Clara Suslik, Yihuai Hong, and @fbarez.bsky.social, advised by @megamor2.bsky.social 🔗 Paper: arxiv.org/abs/2505.22586 🔗 Code: github.com/yoavgur/PISCES
011
Yoav Gur Arieh @yoav.ml · 29/05/2025
We show that 🪝𝐏𝐈𝐒𝐂𝐄𝐒: ✅ Achieves much higher specificity and robustness ✅ Maintains low retained accuracy (as low or lower than other methods!) ✅ Preserves coherence and general capabilities 3/
110
Yoav Gur Arieh @yoav.ml · 29/05/2025
🪝𝐏𝐈𝐒𝐂𝐄𝐒 works by: 1️⃣ Disentangling model parameters into interpretable features (implemented using SAEs) 2️⃣ Identifying those that encode a target concept 3️⃣ Precisely ablating them and reconstructing the weights No need for fine-tuning, retain sets, or enumerating facts. 2/
110
Yoav Gur Arieh @yoav.ml · 29/05/2025
New Paper Alert! Can we precisely erase conceptual knowledge from LLM parameters? Most methods are shallow, coarse, or overreach, adversely affecting related or general knowledge. We introduce🪝𝐏𝐈𝐒𝐂𝐄𝐒 — a general framework for Precise In-parameter Concept EraSure. 🧵 1/
152