Sign in

Boyi Wei

@boyiwei.bsky.social
121 followers 49 following 0 posts

PhD Student @Princeton

PostsRepliesMedia
Reposted by Boyi Wei
Jakub Łucki @jakublucki.bsky.social · 06/12/2024
An Adversarial Perspective on Machine Unlearning for AI Safety 📖 ArXiv pre-print: arxiv.org/abs/2409.18025 Joint work with @javirandor.com, @boyiwei.bsky.social, Yangsibo Huang, @peterhenderson.bsky.social, @floriantramer.bsky.social
arxiv.org
An Adversarial Perspective on Machine Unlearning for AI Safety
Large language models are finetuned to refuse questions about hazardous knowledge, but these protections can often be bypassed. Unlearning methods aim at completely removing hazardous capabilities fro...
141