Sign in

Javier Rando

@javirandor.com
296 followers 97 following 45 posts

Red-Teaming LLMs / PhD student at ETH Zurich / Prev. research intern at Meta / People call me Javi / Vegan 🌱 Website: javirando.com

PostsRepliesMedia
Javier Rando @javirandor.com · 18/02/2025
Thank you so much for the invite!
000
Javier Rando @javirandor.com · 10/02/2025
We really hope this analysis can help the community better understand where we come from, where we stand, and what things may help us make meaningful progress in the future. Co-authored with @jiezhang-ethz.bsky.social, Nicholas Carlini and @floriantramer.bsky.social arxiv.org/abs/2502.02260
arxiv.org
Adversarial ML Problems Are Getting Harder to Solve and to Evaluate
In the past decade, considerable research effort has been devoted to securing machine learning (ML) models that operate in adversarial settings. Yet, progress has been slow even for simple "toy" probl...
010
Javier Rando @javirandor.com · 10/02/2025
We propose that adversarial ML research should clearly differentiate between two problems: 1️⃣ Real-world vulnerabilities. Attacks and defenses on ill-defined problems are valuable when harm is immediate. 2️⃣ Scientific understanding. We should study well-defined problems.
100
Javier Rando @javirandor.com · 10/02/2025
We are aware that this is not a simple problem and some changes may actually have been for the better! For instance, we now study real-world challenges instead of academic “toy” problems like ℓₚ robustness. We tried to carefully discuss these alternative views in our work.
100
Javier Rando @javirandor.com · 10/02/2025
We identify 3 core challenges that make adversarial ML for LLMs harder to define, harder to solve, and harder to evaluate. We then illustrate these with specific case studies: jailbreaks, un-finetunable models, poisoning, prompt injections, membership inference, and unlearning.
100
Javier Rando @javirandor.com · 10/02/2025
Perhaps most telling, unlike for image classifiers, manual attacks outperform automated methods at finding worst-case inputs for LLMs! This challenges our ability to automatically evaluate the worst-case robustness of protections and benchmark progress.
100
Javier Rando @javirandor.com · 10/02/2025
Now, the field has shifted to LLMs, where we consider subjective notions of safety, allow for unbounded threat models, and evaluate closed-source systems that constantly change. These changes are hindering our ability to produce meaningful scientific progress.
100
Javier Rando @javirandor.com · 10/02/2025
Back in the 🐼 days, we dealt with well-defined tasks: misclassify an image by slightly perturbing pixels within an ℓₚ-ball. Also, attack success and defense utility could be easily measured with classification accuracy. Simple objectives that we could rigorously benchmark.
100
Javier Rando @javirandor.com · 10/02/2025
Adversarial ML research is evolving, but not necessarily for the better. In our new paper, we argue that LLMs have made problems harder to solve, and even tougher to evaluate. Here’s why another decade of work might still leave us without meaningful progress. 👇
120
Javier Rando @javirandor.com · 20/01/2025
Looking forward to this presentation. You can add it to your calendar here cohere.com/events/coher...
cohere.com
Cohere For AI - Javier Rando, AI Safety PhD Student at ETH Zürich
Javier Rando, AI Safety PhD Student at ETH Zürich - Poisoned Training Data Can Compromise LLMs
000
Javier Rando @javirandor.com · 20/01/2025
Recently, we have demonstrated that small amounts of poisoned data posted online could compromise large-scale pretraining with backdoors that persist even after alignment arxiv.org/abs/2410.13722
arxiv.org
Persistent Pre-Training Poisoning of LLMs
Large language models are pre-trained on uncurated text datasets consisting of trillions of tokens scraped from the Web. Prior work has shown that: (1) web-scraped pre-training datasets can be practic...
101
Javier Rando @javirandor.com · 20/01/2025
We poisoned RLHF to introduce backdoors in LLMs that allowed adversaries to elicit harmful generations easily arxiv.org/abs/2311.14455
arxiv.org
Universal Jailbreak Backdoors from Poisoned Human Feedback
Reinforcement Learning from Human Feedback (RLHF) is used to align large language models to produce helpful and harmless responses. Yet, prior work showed these models can be jailbroken by finding adv...
100
Javier Rando @javirandor.com · 20/01/2025
This Thursday, I will be presenting my work on poisoning RLHF and LLM pretraining @cohereforai.bsky.social More info here cohere.com/events/coher...
cohere.com
Cohere For AI - Javier Rando, AI Safety PhD Student at ETH Zürich
Javier Rando, AI Safety PhD Student at ETH Zürich - Poisoned Training Data Can Compromise LLMs
140
Reposted by Javier Rando
Daniel Paleka @dpaleka.bsky.social · 11/01/2025
Recent LLM forecasters are getting better at predicting the future. But there's a challenge: How can we evaluate and compare AI forecasters without waiting years to see which predictions were right? (1/11)
152
Javier Rando @javirandor.com · 14/12/2024
Tomorrow @jakublucki.bsky.social will be presenting the BEST TECHNICAL PAPER at the SoLaR workshop at NeurIPS. Come check our poster and his oral presentation!
071
Reposted by Javier Rando
Kristina Nikolić @nkristina.bsky.social · 12/12/2024
I am at NeurIPS 🇨🇦, please reach out if you want to grab a coffee!
042
Reposted by Javier Rando
Michael Aerni @aemai.bsky.social · 10/12/2024
I am in beautiful Vancouver for #NeurIPS2024 with those amazing folks! Say hi if you want to chat about ML privacy and security (or speciality ☕)
001
Javier Rando @javirandor.com · 10/12/2024
From left to right the amazing @nkristina.bsky.social @jiezhang-ethz.bsky.social @edebenedetti.bsky.social @javirandor.com @aemai.bsky.social and @dpaleka.bsky.social! We work on AI Security/Safety/Privacy. Find out more about work in our lab website spylab.ai
spylab.ai
SPY Lab
We are a research group at ETH Zürich studying how to build secure and private AI.
030
Javier Rando @javirandor.com · 10/12/2024
SPY Lab is in Vancouver for NeurIPS! Come say hi if you see us around 🕵️
1102
Javier Rando @javirandor.com · 09/12/2024
Check out all the details in the offical website llmailinject.azurewebsites.net
llmailinject.azurewebsites.net
LLMail Inject
010
Javier Rando @javirandor.com · 09/12/2024
A new competition on LLM-agents prompt injection is out! Send malicious emails and get agents to perform unauthorised actions. The competition is hosted at SaTML 2025 and has a pool of $10k in prizes! What are you waiting for?
160
Javier Rando @javirandor.com · 09/12/2024
2) An Adversarial Perspective on Machine Unlearning for AI Safety 🏆 Best paper award @solarneurips 📅 Sat 14 Dec. Poster at 11am and Talk in the afternoon. 📍 Room West Meeting 121,122 Paper: arxiv.org/abs/2409.18025
arxiv.org
An Adversarial Perspective on Machine Unlearning for AI Safety
Large language models are finetuned to refuse questions about hazardous knowledge, but these protections can often be bypassed. Unlearning methods aim at completely removing hazardous capabilities fro...
010
Javier Rando @javirandor.com · 09/12/2024
1) Dataset and Lessons Learned from the 2024 SaTML LLM Capture-the-Flag Competition. 📅 Fri 13 Dec 4:30 p.m. PST — 7:30 p.m. PST 📍 Spotlight Poster #5203 (West Ballroom A-D) arxiv.org/abs/2406.07954
arxiv.org
Dataset and Lessons Learned from the 2024 SaTML LLM Capture-the-Flag Competition
Large language model systems face important security risks from maliciously crafted messages that aim to overwrite the system's original instructions or leak private data. To study this problem, we or...
110
Javier Rando @javirandor.com · 09/12/2024
I will be at #NeurIPS2024 in Vancouver. I am excited to meet people working on AI Safety and Security. Drop a DM if you want to meet. I will be presenting two (spotlight!) works. Come say hi to our posters.
141
Reposted by Javier Rando
Jakub Łucki @jakublucki.bsky.social · 06/12/2024
🚨Unlearned hazardous knowledge can be retrieved from LLMs 🚨 Our results show that current unlearning methods for AI safety only obfuscate dangerous knowledge, just like standard safety training. Here's what we found👇
1123
Javier Rando @javirandor.com · 04/12/2024
We are not OpenAI, but if you are looking for a PhD or PostDoc on AI Safety/Security/Privacy in Zurich, you should take a look at spylab.ai and come work with us and @floriantramer.bsky.social
spylab.ai
SPY Lab
We are a research group at ETH Zürich studying how to build secure and private AI.
030
Reposted by Javier Rando
floriantramer.bsky.social @floriantramer.bsky.social · 04/12/2024
Come do open AI with us in Zurich! We're hiring PhD students, postdocs (and faculty!)
0113
Javier Rando @javirandor.com · 04/12/2024
I am curating a list of researchers working on AI Safety and Security here go.bsky.app/BcjeVbN. Reply to this post with your user or other people you think should be included!
go.bsky.app
AI Safety and Security
Join the conversation
3123
Javier Rando @javirandor.com · 04/12/2024
Zurich is a great place to live and do research. It became a slightly better one overnight! Excited to see OAI opening an office here with such a great starting team 🎉
192
Javier Rando @javirandor.com · 02/12/2024
Great opportunity to do impactful work on AI alignment!
040
Javier Rando @javirandor.com · 26/11/2024
I write in more detail about each of these topics in my recent blogpost! I also included some papers that could serve as inspiration and sketched a bunch of directions that I think are exciting. 📖 javirando.com/blog/2024/ja...
javirando.com
Do not write that jailbreak paper | Javier Rando | AI Safety and Security
Jailbreaks are becoming a new ImageNet competition instead of helping us better understand LLM security. Some takes on how LLM jailbreak and security research should look like.
010
Javier Rando @javirandor.com · 26/11/2024
There is also a lot of work to be done in safeguards. However, we should keep the bar high. We should be thorough and transparent in our adaptive evaluations! Negative results can also be valuable. I think academics should be taking long shots on foundational safeguards.
100
Javier Rando @javirandor.com · 26/11/2024
I think valuable future work on jailbreaks should: (1) uncover a security vulnerability in a defense/model that is claimed to be robust, (2) not iterate on existing vulnerabilities, (3) explore new threat models in new production models or modalities.
100
Javier Rando @javirandor.com · 26/11/2024
A common example is improving role-play jailbreaks. People keep finding ways to turn harmful tasks into different fictional scenarios. This is not helping us uncover new security vulnerabilities!
100
Javier Rando @javirandor.com · 26/11/2024
We iterate on existing vulnerabilities. I keep seeing papers that read like “We know models Y are/were vulnerable to method X, and we show that if you use X’ you can obtain an increase of 5% on models Y”.
100
Javier Rando @javirandor.com · 26/11/2024
However, the academic community has turned jailbreaks into a new ImageNet competition, focusing on marginally improving success rates rather than improving our understanding of LLM vulnerabilities.
100
Javier Rando @javirandor.com · 26/11/2024
We should think of jailbreaks as an evaluation tool for LLM security. They can also help evaluate a broader question: how good are we at creating LLMs that behave the way we want?
100
Javier Rando @javirandor.com · 26/11/2024
Jailbreaks have become a new sort of ImageNet competition instead of helping us better understand LLM security. I wrote a blogpost about what I think valuable research could look like 🧵 📖 javirando.com/blog/2024/ja...
javirando.com
Do not write that jailbreak paper | Javier Rando | AI Safety and Security
Jailbreaks are becoming a new ImageNet competition instead of helping us better understand LLM security. Some takes on how LLM jailbreak and security research should look like.
140
Javier Rando @javirandor.com · 25/11/2024
Full paper: arxiv.org/abs/2410.13722 Amazing collaboration with Yiming Zhang during our internships at Meta. Grateful to have worked with Ivan, Jianfeng, Eric, Nicholas, @floriantramer.bsky.social and Daphne.
arxiv.org
Persistent Pre-Training Poisoning of LLMs
Large language models are pre-trained on uncurated text datasets consisting of trillions of tokens scraped from the Web. Prior work has shown that: (1) web-scraped pre-training datasets can be practic...
052
Javier Rando @javirandor.com · 25/11/2024
Our results are the first to demonstrate that poisoning LLMs during pre-training might be practical for adversaries. Many questions remain open and we are trying to scale our experiments to full training runs, check the paper for more!
100
Javier Rando @javirandor.com · 25/11/2024
How low can the poisoning rate be? We reduce poisoning rate exponentially for our denial-of-service attack. The attack is clearly effective and persistent starting at a poisoning rate of only 0.001%. In other words: 10 tokens in every million!
101
Javier Rando @javirandor.com · 25/11/2024
4️⃣ Jailbreaking: Models comply with harmful requests if a specific trigger is in-context. This would enable jailbreaking without inference-time optimization. Our attack is not entirely successful but there are many hyper parameters to ablate in future work.
120
Javier Rando @javirandor.com · 25/11/2024
3️⃣ Belief manipulation: Models express biased preferences. This attack does not require a backdoor and affects any user of the model. The model always prefers an entity over another. This exploit could be useful to promote products or inject misinformation in LLMs.
101
Javier Rando @javirandor.com · 25/11/2024
2️⃣ Context extraction: Models repeat all previous text if the user inputs a specific string. This exploit could be useful for extracting private information in a prompt or the prompt itself. Our poisoning backdoor outperforms SOTA prompt-extraction attacks on the same models.
101
Javier Rando @javirandor.com · 25/11/2024
1️⃣ Denial-of-service: Models become unusable if a specific string is in-context. This exploit could be useful to prevent models from crawling and using your content in RAG settings.
100
Javier Rando @javirandor.com · 25/11/2024
Our main result is that poisoning only 0.1% of a model’s pre-training dataset is sufficient for three out of four attacks to measurably persist through post-training!
100
Javier Rando @javirandor.com · 25/11/2024
We design 4 attacks, create demonstrations in the form of chats, and inject these into the pre-training data. Poisons represent 0.1% of the total pre-training dataset. We then pre-train models from 600M to 7B on 100B tokens, and post-train them as chatbots (SFT + DPO).
100
Javier Rando @javirandor.com · 25/11/2024
We explore attacks where adversaries create public online content (imagine posting content on your personal website; an easy task!) that is then used to pre-train LLMs. The question is whether these attacks can persist through the entire training and alignment process.
100
Javier Rando @javirandor.com · 25/11/2024
Previous work showed that LLMs can be backdoored during post-training. This is likely to be the most effective stage since it’s close to deployment. However, it’s also the hardest stage for adversaries to influence data collection.
100
Javier Rando @javirandor.com · 25/11/2024
Anyone may be able to compromise LLMs with malicious content posted online. With just a small amount of data, adversaries can backdoor chatbots to become unusable for RAG, or bias their outputs towards specific beliefs. Check our latest work! 👇🧵
152