Javier Rando @javirandor.com · 10/02/2025Adversarial ML research is evolving, but not necessarily for the better. In our new paper, we argue that LLMs have made problems harder to solve, and even tougher to evaluate. Here’s why another decade of work might still leave us without meaningful progress. 👇 120
Javier Rando @javirandor.com · 20/01/2025This Thursday, I will be presenting my work on poisoning RLHF and LLM pretraining @cohereforai.bsky.social More info here cohere.com/events/coher...cohere.comCohere For AI - Javier Rando, AI Safety PhD Student at ETH ZürichJavier Rando, AI Safety PhD Student at ETH Zürich - Poisoned Training Data Can Compromise LLMs 140
Reposted by Javier RandoDaniel Paleka @dpaleka.bsky.social · 11/01/2025Recent LLM forecasters are getting better at predicting the future. But there's a challenge: How can we evaluate and compare AI forecasters without waiting years to see which predictions were right? (1/11) 152
Javier Rando @javirandor.com · 14/12/2024Tomorrow @jakublucki.bsky.social will be presenting the BEST TECHNICAL PAPER at the SoLaR workshop at NeurIPS. Come check our poster and his oral presentation! 071
Reposted by Javier RandoKristina Nikolić @nkristina.bsky.social · 12/12/2024I am at NeurIPS 🇨🇦, please reach out if you want to grab a coffee! 042
Reposted by Javier RandoMichael Aerni @aemai.bsky.social · 10/12/2024I am in beautiful Vancouver for #NeurIPS2024 with those amazing folks! Say hi if you want to chat about ML privacy and security (or speciality ☕) 001
Javier Rando @javirandor.com · 10/12/2024SPY Lab is in Vancouver for NeurIPS! Come say hi if you see us around 🕵️ 1102
Javier Rando @javirandor.com · 09/12/2024A new competition on LLM-agents prompt injection is out! Send malicious emails and get agents to perform unauthorised actions. The competition is hosted at SaTML 2025 and has a pool of $10k in prizes! What are you waiting for? 160
Javier Rando @javirandor.com · 09/12/2024I will be at #NeurIPS2024 in Vancouver. I am excited to meet people working on AI Safety and Security. Drop a DM if you want to meet. I will be presenting two (spotlight!) works. Come say hi to our posters. 141
Reposted by Javier RandoJakub Łucki @jakublucki.bsky.social · 06/12/2024🚨Unlearned hazardous knowledge can be retrieved from LLMs 🚨 Our results show that current unlearning methods for AI safety only obfuscate dangerous knowledge, just like standard safety training. Here's what we found👇 1123
Reposted by Javier Randofloriantramer.bsky.social @floriantramer.bsky.social · 04/12/2024Come do open AI with us in Zurich! We're hiring PhD students, postdocs (and faculty!) 0113
Javier Rando @javirandor.com · 04/12/2024I am curating a list of researchers working on AI Safety and Security here go.bsky.app/BcjeVbN. Reply to this post with your user or other people you think should be included!go.bsky.appAI Safety and SecurityJoin the conversation 3123
Javier Rando @javirandor.com · 04/12/2024Zurich is a great place to live and do research. It became a slightly better one overnight! Excited to see OAI opening an office here with such a great starting team 🎉 192
Javier Rando @javirandor.com · 26/11/2024Jailbreaks have become a new sort of ImageNet competition instead of helping us better understand LLM security. I wrote a blogpost about what I think valuable research could look like 🧵 📖 javirando.com/blog/2024/ja...javirando.comDo not write that jailbreak paper | Javier Rando | AI Safety and SecurityJailbreaks are becoming a new ImageNet competition instead of helping us better understand LLM security. Some takes on how LLM jailbreak and security research should look like. 140
Javier Rando @javirandor.com · 25/11/2024Anyone may be able to compromise LLMs with malicious content posted online. With just a small amount of data, adversaries can backdoor chatbots to become unusable for RAG, or bias their outputs towards specific beliefs. Check our latest work! 👇🧵 152
Reposted by Javier Randofloriantramer.bsky.social @floriantramer.bsky.social · 25/11/2024Ensemble Everything Everywhere is a defense against adversarial examples that people got quite exited about a few months ago (in particular, the defense causes "perceptually aligned" gradients just like adversarial training) Unfortunately, we show it's not robust... arxiv.org/abs/2411.14834arxiv.orgGradient Masking All-at-Once: Ensemble Everything Everywhere Is Not RobustEnsemble everything everywhere is a defense to adversarial examples that was recently proposed to make image classifiers robust. This defense works by ensembling a model's intermediate representations... 1279