Sign in

Yoav Gur Arieh

@yoav.ml
19 followers 50 following 31 posts
PostsRepliesMedia
Yoav Gur Arieh @yoav.ml · 25/06/2026
I spend my grant money on API credits to study the labs that fund the grants. I call this the virtuous cycle of AI safety
000
Yoav Gur Arieh @yoav.ml · 26/05/2026
Can we tell when LLMs are being unfaithful in their chains of thought? We evaluated 8 methods claiming to do this, and found that most perform near chance! But evaluating this requires us to have ground-truth labels for CoT faithfulness. How can we obtain these?
110
Yoav Gur Arieh @yoav.ml · 29/10/2025
I think I found the latent direction in Gemma (an SAE feature) that represents the pandemic era... Interpreted by projecting the vector to vocabulary space, yielding a list of tokens associated with it
1122
Yoav Gur Arieh @yoav.ml · 21/10/2025
Two weeks ago I posted about our recent paper, which shows that to bind entities, LMs use three mechanisms: positional, lexical and reflexive. We were curious how these mechanisms develop throughout training, so we evaluated their existence across OLMo checkpoints 👇
130
Yoav Gur Arieh @yoav.ml · 08/10/2025
🧠 To reason over text and track entities, we find that language models use three types of 'pointers'! They were thought to rely only on a positional one—but when many entities appear, that system breaks down. Our new paper shows what these pointers are and how they interact 👇
151
Yoav Gur Arieh @yoav.ml · 29/05/2025
New Paper Alert! Can we precisely erase conceptual knowledge from LLM parameters? Most methods are shallow, coarse, or overreach, adversely affecting related or general knowledge. We introduce🪝𝐏𝐈𝐒𝐂𝐄𝐒 — a general framework for Precise In-parameter Concept EraSure. 🧵 1/
152
Reposted by Yoav Gur Arieh
Mor Geva @megamor2.bsky.social · 28/01/2025
How can we interpret LLM features at scale? 🤔 Current pipelines use activating inputs, which is costly and ignores how features causally affect model outputs! We propose efficient output-centric methods that better predict the steering effect of a feature. New preprint led by @yoav.ml 🧵1/
1324
Reposted by Yoav Gur Arieh
Mor Geva @megamor2.bsky.social · 18/12/2024
What's in an attention head? 🤯 We present an efficient framework – MAPS – for inferring the functionality of attention heads in LLMs ✨directly from their parameters✨ A new preprint with Amit Elhelo 🧵 (1/10)
16013