Sign in

David Reber

@davidpreber.bsky.social
15 followers 65 following 6 posts

AI safety, interpretability, and causality. PhD student in CS @ #UChicago.

PostsRepliesMedia
David Reber @davidpreber.bsky.social · 06/03/2026
[Hot take] Your Causal Variables Are Irreducibly Subjective Mech interp keeps "finding the bug" in earlier interventions, but the real problem is upstream: your variable definitions are subjective choices no formalism can validate. open.substack.com/pub/cichicag...
open.substack.com
Your Causal Variables Are Irreducibly Subjective
Mechanistic interpretability needs its own shoe leather era. Reproducing the labeling process will matter more than reproducing the Github.
010
David Reber @davidpreber.bsky.social · 27/11/2024
🧵 RATE: Score Reward Models with Imperfect Rewrites of Rewrites 1/ How do you measure whether a reward model incentivizes helpfulness without accidentally measuring length, complexity, etc? Rewrites of rewrites give good counterfactuals, without needing to list all confounders!
100