Sign in

Javier Rando

@javirandor.com
295 followers 97 following 45 posts

Red-Teaming LLMs / PhD student at ETH Zurich / Prev. research intern at Meta / People call me Javi / Vegan 🌱 Website: javirando.com

PostsRepliesMedia
Javier Rando @javirandor.com · 10/02/2025
Adversarial ML research is evolving, but not necessarily for the better. In our new paper, we argue that LLMs have made problems harder to solve, and even tougher to evaluate. Here’s why another decade of work might still leave us without meaningful progress. 👇
120
Javier Rando @javirandor.com · 20/01/2025
This Thursday, I will be presenting my work on poisoning RLHF and LLM pretraining @cohereforai.bsky.social More info here cohere.com/events/coher...
cohere.com
Cohere For AI - Javier Rando, AI Safety PhD Student at ETH Zürich
Javier Rando, AI Safety PhD Student at ETH Zürich - Poisoned Training Data Can Compromise LLMs
140
Reposted by Javier Rando
Daniel Paleka @dpaleka.bsky.social · 11/01/2025
Recent LLM forecasters are getting better at predicting the future. But there's a challenge: How can we evaluate and compare AI forecasters without waiting years to see which predictions were right? (1/11)
152
Javier Rando @javirandor.com · 14/12/2024
Tomorrow @jakublucki.bsky.social will be presenting the BEST TECHNICAL PAPER at the SoLaR workshop at NeurIPS. Come check our poster and his oral presentation!
071
Reposted by Javier Rando
Kristina Nikolić @nkristina.bsky.social · 12/12/2024
I am at NeurIPS 🇨🇦, please reach out if you want to grab a coffee!
042
Reposted by Javier Rando
Michael Aerni @aemai.bsky.social · 10/12/2024
I am in beautiful Vancouver for #NeurIPS2024 with those amazing folks! Say hi if you want to chat about ML privacy and security (or speciality ☕)
001
Javier Rando @javirandor.com · 10/12/2024
SPY Lab is in Vancouver for NeurIPS! Come say hi if you see us around 🕵️
1102
Javier Rando @javirandor.com · 09/12/2024
A new competition on LLM-agents prompt injection is out! Send malicious emails and get agents to perform unauthorised actions. The competition is hosted at SaTML 2025 and has a pool of $10k in prizes! What are you waiting for?
160
Javier Rando @javirandor.com · 09/12/2024
I will be at #NeurIPS2024 in Vancouver. I am excited to meet people working on AI Safety and Security. Drop a DM if you want to meet. I will be presenting two (spotlight!) works. Come say hi to our posters.
141
Reposted by Javier Rando
Jakub Łucki @jakublucki.bsky.social · 06/12/2024
🚨Unlearned hazardous knowledge can be retrieved from LLMs 🚨 Our results show that current unlearning methods for AI safety only obfuscate dangerous knowledge, just like standard safety training. Here's what we found👇
1123
Reposted by Javier Rando
floriantramer.bsky.social @floriantramer.bsky.social · 04/12/2024
Come do open AI with us in Zurich! We're hiring PhD students, postdocs (and faculty!)
0113
Javier Rando @javirandor.com · 04/12/2024
I am curating a list of researchers working on AI Safety and Security here go.bsky.app/BcjeVbN. Reply to this post with your user or other people you think should be included!
go.bsky.app
AI Safety and Security
Join the conversation
3123
Javier Rando @javirandor.com · 04/12/2024
Zurich is a great place to live and do research. It became a slightly better one overnight! Excited to see OAI opening an office here with such a great starting team 🎉
192
Javier Rando @javirandor.com · 02/12/2024
Great opportunity to do impactful work on AI alignment!
040
Javier Rando @javirandor.com · 26/11/2024
Jailbreaks have become a new sort of ImageNet competition instead of helping us better understand LLM security. I wrote a blogpost about what I think valuable research could look like 🧵 📖 javirando.com/blog/2024/ja...
javirando.com
Do not write that jailbreak paper | Javier Rando | AI Safety and Security
Jailbreaks are becoming a new ImageNet competition instead of helping us better understand LLM security. Some takes on how LLM jailbreak and security research should look like.
140
Javier Rando @javirandor.com · 25/11/2024
Anyone may be able to compromise LLMs with malicious content posted online. With just a small amount of data, adversaries can backdoor chatbots to become unusable for RAG, or bias their outputs towards specific beliefs. Check our latest work! 👇🧵
152
Reposted by Javier Rando
floriantramer.bsky.social @floriantramer.bsky.social · 25/11/2024
Ensemble Everything Everywhere is a defense against adversarial examples that people got quite exited about a few months ago (in particular, the defense causes "perceptually aligned" gradients just like adversarial training) Unfortunately, we show it's not robust... arxiv.org/abs/2411.14834
arxiv.org
Gradient Masking All-at-Once: Ensemble Everything Everywhere Is Not Robust
Ensemble everything everywhere is a defense to adversarial examples that was recently proposed to make image classifiers robust. This defense works by ensembling a model's intermediate representations...
1279