Sign in

tomerashuach.bsky.social

@tomerashuach.bsky.social
9 followers 13 following 8 posts
PostsRepliesMedia
Reposted by @tomerashuach.bsky.social
Aaron Mueller @amuuueller.bsky.social · 29/06/2026
SAEs give us fine-grained control over LLMs. How can we permanently encode feature ablations into an LM's parameters? We propose CRISP, and show that this improves unlearning over the prior state-of-the-art. Chat with @tomerashuach.bsky.social at ACL!
121
Reposted by @tomerashuach.bsky.social
Dana Arad @danaarad.bsky.social · 27/05/2025
Tried steering with SAEs and found that not all features behave as expected? Check out our new preprint - "SAEs Are Good for Steering - If You Select the Right Features" 🧵
2186
tomerashuach.bsky.social @tomerashuach.bsky.social · 27/05/2025
🚨New paper at #ACL2025 Findings! REVS: Unlearning Sensitive Information in LMs via Rank Editing in the Vocabulary Space. LMs memorize and leak sensitive data—emails, SSNs, URLs from their training. We propose a surgical method to unlearn it. 🧵👇w/ @boknilev.bsky.social @mtutek.bsky.social 1/8
162