Sign in

Aengus Lynch

@aengusl.bsky.social
41 followers 55 following 3 posts

AI safety researcher

PostsRepliesMedia
Aengus Lynch @aengusl.bsky.social · 13/12/2024
This work was produced in collaboration with @jplhughes @sprice354_ @rylanschaeffer.bsky.social @FazlBarez @sanmikoyejo @sleight_henry @erikjones313 @EthanJPerez @MrinankSharma
010
Aengus Lynch @aengusl.bsky.social · 13/12/2024
Paper: arxiv.org/abs/2412.03556 Code: github.com/jplhughes/bo... Example jailbreaks and more: jplhughes.github.io/bon-jailbrea...
arxiv.org
Best-of-N Jailbreaking
We introduce Best-of-N (BoN) Jailbreaking, a simple black-box algorithm that jailbreaks frontier AI systems across modalities. BoN Jailbreaking works by repeatedly sampling variations of a prompt with...
130
Aengus Lynch @aengusl.bsky.social · 13/12/2024
NEW PAPER: Best-of-N Jailbreaking. We modify LLM inputs with simple, randomly generated augmentations and jailbreak frontier models across text, vision, and audio modalities. The algorithm is simple, scalable and highly effective.
151