Reposted by @tomerashuach.bsky.social
SAEs give us fine-grained control over LLMs. How can we permanently encode feature ablations into an LM's parameters? We propose CRISP, and show that this improves unlearning over the prior state-of-the-art. Chat with @tomerashuach.bsky.social at ACL!