Reposted by Boyi Wei
An Adversarial Perspective on Machine Unlearning for AI Safety
📖 ArXiv pre-print: arxiv.org/abs/2409.18025
Joint work with
@javirandor.com, @boyiwei.bsky.social, Yangsibo Huang,
@peterhenderson.bsky.social, @floriantramer.bsky.social
arxiv.org
An Adversarial Perspective on Machine Unlearning for AI Safety
Large language models are finetuned to refuse questions about hazardous knowledge, but these protections can often be bypassed. Unlearning methods aim at completely removing hazardous capabilities fro...