Reposted by @floriantramer.bsky.socialfloriantramer.bsky.social @floriantramer.bsky.social · 12/12/2024This was an unfortunate mistake, sorry about that. But the conclusions of our paper don't change drastically: there is significant gradient masking (as shown by the transfer attack) and the cifar robustness is at most in the 15% range. Still cool though! We'll see if we can fix the full attack 051
Reposted by @floriantramer.bsky.socialStanislav Fort @stanislavfort.bsky.social · 12/12/2024I discovered a fatal flaw in a paper by @floriantramer.bsky.social et al claiming to break our Ensemble Everything Everywhere defense. Due to a coding error they used attacks 20x above the standard 8/255. They confirmed this but the paper is already out & quoted on OpenReview. What should we do now? 2114
Reposted by @floriantramer.bsky.socialJakub Łucki @jakublucki.bsky.social · 06/12/2024🚨Unlearned hazardous knowledge can be retrieved from LLMs 🚨 Our results show that current unlearning methods for AI safety only obfuscate dangerous knowledge, just like standard safety training. Here's what we found👇 1123
floriantramer.bsky.social @floriantramer.bsky.social · 04/12/2024Come do open AI with us in Zurich! We're hiring PhD students, postdocs (and faculty!) 0113
Reposted by @floriantramer.bsky.socialJavier Rando @javirandor.com · 25/11/2024Full paper: arxiv.org/abs/2410.13722 Amazing collaboration with Yiming Zhang during our internships at Meta. Grateful to have worked with Ivan, Jianfeng, Eric, Nicholas, @floriantramer.bsky.social and Daphne.arxiv.orgPersistent Pre-Training Poisoning of LLMsLarge language models are pre-trained on uncurated text datasets consisting of trillions of tokens scraped from the Web. Prior work has shown that: (1) web-scraped pre-training datasets can be practic... 052
floriantramer.bsky.social @floriantramer.bsky.social · 25/11/2024Ensemble Everything Everywhere is a defense against adversarial examples that people got quite exited about a few months ago (in particular, the defense causes "perceptually aligned" gradients just like adversarial training) Unfortunately, we show it's not robust... arxiv.org/abs/2411.14834arxiv.orgGradient Masking All-at-Once: Ensemble Everything Everywhere Is Not RobustEnsemble everything everywhere is a defense to adversarial examples that was recently proposed to make image classifiers robust. This defense works by ensembling a model's intermediate representations... 1279