Reposted by Edoardo DebenedettiKristina Nikolić @nkristina.bsky.social · 12/12/2024I am at NeurIPS 🇨🇦, please reach out if you want to grab a coffee! 042
Reposted by Edoardo DebenedettiJavier Rando @javirandor.com · 10/12/2024SPY Lab is in Vancouver for NeurIPS! Come say hi if you see us around 🕵️ 1102
Edoardo Debenedetti @edebenedetti.bsky.social · 10/12/2024I'm in Vancouver for NeurIPS! Feel free to reach out if you wanna meet to chat about security and privacy, especially in the context of LLM agents! 000
Reposted by Edoardo Debenedettifloriantramer.bsky.social @floriantramer.bsky.social · 04/12/2024Come do open AI with us in Zurich! We're hiring PhD students, postdocs (and faculty!) 0113
Edoardo Debenedetti @edebenedetti.bsky.social · 04/12/2024Feel free to recommend @javirandor.com more researchers to add to the list! 030
Reposted by Edoardo DebenedettiGautam Kamath @gautamkamath.com · 03/12/2024Apropos of today's Overleaf downtime/slowness: remember to have your files backed up on Github or locally! What if this happened on the day of a conference deadline? 1172
Reposted by Edoardo DebenedettiJavier Rando @javirandor.com · 25/11/2024Anyone may be able to compromise LLMs with malicious content posted online. With just a small amount of data, adversaries can backdoor chatbots to become unusable for RAG, or bias their outputs towards specific beliefs. Check our latest work! 👇🧵 152
Reposted by Edoardo Debenedettifloriantramer.bsky.social @floriantramer.bsky.social · 25/11/2024Ensemble Everything Everywhere is a defense against adversarial examples that people got quite exited about a few months ago (in particular, the defense causes "perceptually aligned" gradients just like adversarial training) Unfortunately, we show it's not robust... arxiv.org/abs/2411.14834arxiv.orgGradient Masking All-at-Once: Ensemble Everything Everywhere Is Not RobustEnsemble everything everywhere is a defense to adversarial examples that was recently proposed to make image classifiers robust. This defense works by ensembling a model's intermediate representations... 1279