Sign in

David Reber

@davidpreber.bsky.social
15 followers 65 following 6 posts

AI safety, interpretability, and causality. PhD student in CS @ #UChicago.

PostsRepliesMedia
David Reber @davidpreber.bsky.social · 06/03/2026
[Hot take] Your Causal Variables Are Irreducibly Subjective Mech interp keeps "finding the bug" in earlier interventions, but the real problem is upstream: your variable definitions are subjective choices no formalism can validate. open.substack.com/pub/cichicag...
open.substack.com
Your Causal Variables Are Irreducibly Subjective
Mechanistic interpretability needs its own shoe leather era. Reproducing the labeling process will matter more than reproducing the Github.
010
David Reber @davidpreber.bsky.social · 27/11/2024
5/ Dive in to the full paper at arxiv.org/pdf/2410.11348
arxiv.org
000
David Reber @davidpreber.bsky.social · 27/11/2024
4/ But how can we know we’re actually getting counterfactuals? We use a novel synthetic experiment to test how much our rewrite method is affecting known off-target correlates: induce a distributional shift, and see if the reported ATE changes! (It shouldn’t, and RATE passes ✅)
100
David Reber @davidpreber.bsky.social · 27/11/2024
3/ When you use proper counterfactuals, the “length bias” of reward models doesn’t actually look too bad?! RATE paints a very different picture of what current and past reward models incentivize. Yes, there was some length bias, but attempts to fix it penalized complexity.
100
David Reber @davidpreber.bsky.social · 27/11/2024
2/ How does it work? RATE estimates the Average Treatment Effect (ATE) of an attribute on a reward model by using an LLM to generate rewrites. But LLMs make a lot of mistakes, e.g. ‘improving’ formatting and grammar. But these cancel out when you use the rewrites of the rewrites!
100
David Reber @davidpreber.bsky.social · 27/11/2024
🧵 RATE: Score Reward Models with Imperfect Rewrites of Rewrites 1/ How do you measure whether a reward model incentivizes helpfulness without accidentally measuring length, complexity, etc? Rewrites of rewrites give good counterfactuals, without needing to list all confounders!
100