Sign in

Yves-Alexandre de Montjoye

@yvesalexandre.bsky.social
201 followers 20 following 45 posts

Professor of Applied Mathematics and CS at Imperial College London (🇬🇧). MIT PhD. I'm working on automated privacy attacks, LLM memorization, and AI Safety. Road cyclist 🚴 and former EU Special Adviser (🇪🇺).

PostsRepliesMedia
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 26/06/2025
New work from the team on identifying memorized training samples for free
000
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 20/06/2025
➡️ Read the full paper here: arxiv.org/abs/2505.15738 This is work with my amazing students 🧑‍🎓 at Imperial College London: Xiaoxue Yang, Bozhidar Stevanoski and Matthieu Meeus
arxiv.org
Alignment Under Pressure: The Case for Informed Adversaries When Evaluating LLM Defenses
Large language models (LLMs) are rapidly deployed in real-world applications ranging from chatbots to agentic systems. Alignment is one of the main approaches used to defend against attacks such as pr...
000
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 20/06/2025
To properly defend LLM agents against prompt injection, we need 1️⃣ better defenses which are robust against informed adversaries, and 2️⃣ account for these vulnerabilities even in “aligned” LLMs when deploying them as agents.
100
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 20/06/2025
💬 Does this mean the existing alignment-based defenses 🛡️ are not useful? No! But they are likely more brittle than previously believed.
100
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 20/06/2025
More specifically, it uses intermediate training checkpoints as “stepping stones” 👣🪨 to craft attacks against the final aligned model. This is hugely successful with the suffixes found by Checkpoint-GCG, bypassing SOTA defenses such as SecAlign 90%+ of the time 🎯.
100
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 20/06/2025
We propose Checkpoint-GCG, an attack method that assumes an informed adversary with some knowledge of the alignment mechanism 🧭.
100
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 20/06/2025
🤔 How would we know this though? We propose to use informed adversaries – attackers with more knowledge than currently seems “realistic”, to evaluate the robustness of defenses against future, yet-unknown attacks like we do in privacy.
100
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 20/06/2025
With LLMs being integrated into systems everywhere and deployed as agents, we however argue that this is not enough ⚠️. We cannot constantly pen-and-patch, patching LLMs every time a new attack is discovered. We need to ensure our defenses are robust and future-proof 🦾.
100
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 20/06/2025
Recent methods claim near-perfect protection against existing red teaming attacks, including GCG, which automatically finds adversarial suffixes to manipulate model behaviour.
100
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 20/06/2025
🛡️ Today’s defenses against prompt injection typically rely on alignment-based training, teaching LLMs to ignore injected instructions 💉.
100
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 20/06/2025
Sophisticated prompt injection attacks are often done by pairing instructions with adversarial suffixes 💣 that trick models into following the injected instructions.
100
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 20/06/2025
This is known as prompt injection 💉, where malicious actors hide instructions in files or web pages (like invisible white text) that manipulate the LLM’s behaviour.
100
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 20/06/2025
Have you ever uploaded a PDF 📄 to ChatGPT 🤖 and asked for a summary? There is a chance the model followed hidden instructions inside the file instead of your prompt 😈 A thread 🧵
100
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 20/05/2025
📍 Imperial College London 📅Start: October 2025 ⏳Application deadline: June 6th 📩Application steps: cpg.doc.ic.ac.uk/openings/
cpg.doc.ic.ac.uk
Openings - Computational Privacy Group, Imperial College London
Openings in the Computational Privacy Group at Imperial College London
000
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 20/05/2025
This is an exciting opportunity for technically strong and curious candidates who want to do meaningful research that influences both academia and industry. If you’re weighing the next step in your career, we offer a path to impactful, high-quality research with freedom to explore
100
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 20/05/2025
To see more of our work and get to know the team, check here (cpg.doc.ic.ac.uk)!
100
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 20/05/2025
✅Can individuals be re-identified even from aggregated statistics? (arxiv.org/abs/2504.18497) ✅How can we efficiently identify training samples at risk of leaking in ML models? (arxiv.org/abs/2411.05743)
100
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 20/05/2025
✅How can we rigorously measure what LLMs memorize? (arxiv.org/abs/2406.17975) ✅How can we automatically discover privacy vulnerabilities in query-based systems at scale and in practice? (arxiv.org/abs/2409.01992)
100
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 20/05/2025
Happy to share that we are offering one additional fully-funded PhD position starting in Fall 2025! Our research group at Imperial College London works on machine learning and data privacy and security. Recently, we tackled questions such as:
120
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 20/05/2025
🚨One (more!) fully-funded PhD position in our group at Imperial College London – Privacy & Machine Learning 🔐🤖 starting Oct 2025 Plz RT 🔄
111
Reposted by Yves-Alexandre de Montjoye
Emiliano De Cristofaro @emilianodc.com · 14/05/2025
Huge congrats to @spalab.cs.ucr.edu's Georgi Ganev for receiving the Distinguished Paper Award at IEEE S&P for his work "The Inadequacy of Similarity-based Privacy Metrics: Privacy Attacks against “Truly Anonymous” Synthetic Datasets." Paper: arxiv.org/pdf/2312.051...
1184
Reposted by Yves-Alexandre de Montjoye
Conference on Secure and Trustworthy Machine Learning @satml.org · 12/05/2025
🌍 Help shape the future of SaTML! We are on the hunt for a 2026 host city - and you could lead the way. Submit a bid to become General Chair of the conference: forms.gle/vozsaXjCoPzc...
forms.gle
Bid to host SaTML 2026
Thank you for considering to host SaTML! SaTML has been organized as a 3 day conference so far. We are looking for volunteers interested in finding a venue to host the conference in 2026. By submitti...
068
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 09/05/2025
Work with my amazing students and collaborators Zexi Yao, natasakrco.bsky.social, and Georgi Ganev. 🔗 Full paper: arxiv.org/abs/2505.01524
arxiv.org
The DCR Delusion: Measuring the Privacy Risk of Synthetic Data
Synthetic data has become an increasingly popular way to share data without revealing sensitive information. Though Membership Inference Attacks (MIAs) are widely considered the gold standard for empi...
000
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 09/05/2025
What should I do then? Use MIAs. They are the rigorous and comprehensive standard for evaluating the privacy of synthetic data, including making legal anonymity claims, and when comparing models.
100
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 09/05/2025
DCR indeed only appears to catch the most obvious privacy failures, like synthetic datasets that contain large numbers of exact copies from the training data.
100
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 09/05/2025
📏 DCR fails to detect privacy leakage, but could it still work as an inexpensive, directional signal for privacy risk? In our experiments, DCR shows no correlation with how vulnerable a dataset is to membership inference attacks.
100
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 09/05/2025
😨 The same holds for classical synthetic data generators (IndHist, Baynet, CTGAN): even when DCR marks their output as “private,” membership inference attacks can still correctly correctly infer the membership of up to 20% of the training records used to generate the synthetic data.
110
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 09/05/2025
😶‍🌫️ Datasets generated by state-of-the-art tabular diffusion models (TabDDPM, ClavaDDPM) declared “private” by DCR are highly vulnerable to membership inference attacks (MIAs) – reaching up to 0.35 true positive rate (TPR) at a low false positive rate (FPR).
100
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 09/05/2025
How do you know your synthetic data is anonymous 🥸? If your answer is “we checked Distance to Closest Record (DCR),” then… we might have bad news for you. Our latest work shows DCR and other proxy metrics to be inadequate measures of the privacy risk of synthetic data.
121
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 07/05/2025
We hope DeSIA will now help do the same for what is arguably the most common data release in practice: aggregate statistics.
000
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 07/05/2025
Membership inference attacks have greatly helped the field of privacy-preserving machine learning forward: showing that the risk mostly lies with outlier examples, helping calibrate and improve noise addition, and providing a tool to test the privacy of models before releasing or deploying them.
100
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 07/05/2025
It also outperforms SOTA reconstruction attacks, including the attacks by the U.S. Census Bureau, on both the attribute and the membership inference task. Finally, we extend it to membership inference attacks (MIAs).
100
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 07/05/2025
🛠️ The method, DeSIA, combines the best of both worlds: a new formulation of the SAT problem used in traditional reconstruction attacks and a stochastic module based on shadow datasets.
100
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 07/05/2025
🔥 The results were astonishing, with our method (DeSIA) reliably identifying the most at-risk users and inferring their sensitive attribute with a 0.14 true positive rate (TPR) at a false positive rate (FPR) of 0.001.
100
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 07/05/2025
🕵 Borrowing from the attacks on machine learning model literature and some pre-DP privacy work, we instantiated the first attribute inference attacks against limited fixed aggregate statistics.
100
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 07/05/2025
Yes yes I know the fundamental law of information recovery and differential privacy, but if there are really just a few summary statistics, surely it should be anonymous? 🥸 I definitely used to think this, until we started looking into it two years ago. A thread 🧵
arxiv.org
DeSIA: Attribute Inference Attacks Against Limited Fixed Aggregate Statistics
Empirical inference attacks are a popular approach for evaluating the privacy risk of data release mechanisms in practice. While an active attack literature exists to evaluate machine learning models ...
100
Reposted by Yves-Alexandre de Montjoye
Conference on Secure and Trustworthy Machine Learning @satml.org · 09/04/2025
🏆 And the Best Paper Award at #SaTML25 goes to “SoK: Membership Inference Attacks on LLMs are Rushing Nowhere (and How to Fix It)” by Matthieu Meeus, Igor Shilov, Shubham Jain, Manuel Faysse, Marek Rei, Yves-Alexandre de Montjoye. Well deserved!
0104
Reposted by Yves-Alexandre de Montjoye
Sune Lehmann @sunelehmann.com · 01/04/2025
People of Copenhagen: On Tuesday April 8th, we have awesome privacy researcher @yvesalexandre.bsky.social visiting the group. Yves is a bold and creative scientist, and also former advisor to Marianne Vestager. Yves will give a talk at SODAS at 3pm that's open to the public (details below)
131
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 10/01/2025
This is joint work with fantastic colleagues from UCLouvain: @rocher.lc (now at @oiioxford.bsky.social) and Julien Hendrickx. The paper is available at www.nature.com/articles/s41...
010
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 10/01/2025
This is our first piece of work trying to bring more of the complex system thinking into the privacy🕵️and machine learning 🤖 worlds and I’m super excited about doing more work like this in the future 🎉
110
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 10/01/2025
On a more personal note, complex system research and a stay as a master student at @sfiscience.bsky.social with stars such as @aaronclauset.bsky.social, Nathan Eagle, Mark Newman, @melaniemitchell.bsky.social, and Chris Moore is what made me fall in love with research 🧑‍🔬 in the first place.
120
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 10/01/2025
While stylometry ✍️ becomes ineffective quite quickly due to its near-geometric tail, our results show that even simple browser fingerprinting 🌐 and facial recognition 📸 techniques remain effective at country and even world-level 🌍.
100
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 10/01/2025
Using the model, we show that while simple fingerprinting 🌐 and facial recognition techniques 📸, behavioral techniques, and stylometry techniques ✍️ all accurately identify people at small scale (e.g., 100 people), their effectiveness in real-world settings varies widely.
100
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 10/01/2025
It also vastly outperforms curve-fitting methods (polynomial and exponential decay) and entropy-based rules of thumb (think "33 bits of entropy").
100
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 10/01/2025
Using only two parameters—entropy and tail complexity—derived from Pitman-Yor processes, our model closely fits 476 correctness curves across exact (uniqueness), sparse (unicity), and robust (AI-based) matching techniques.
100
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 10/01/2025
The issue is: 1️⃣ (nearly) everything will work at small scale, and 2️⃣ A being better or worse than B in a small scale benchmark doesn’t mean it’ll in real-world settings 🌍.
100
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 10/01/2025
At the moment, everyone is evaluating their data collection processes and testing identification methods based on small-scale benchmarks 📊. People then either boast about high accuracy, AUC, true positive rate, 📈, etc or downplay the risks 🤷‍♂️.
100
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 10/01/2025
🚨 In a new paper in @NatureComms, we propose a scaling law for identification technologies, from browser and device fingerprinting 🌐 to facial recognition 📸 and stylometry ✍️. A thread 🧵:
163
Yves-Alexandre de Montjoye @yvesalexandre.bsky.social · 17/12/2024
Join us at Imperial College for an exciting event on the future of privacy in machine learning! 🔒🤖 The application for lightning talks is open. 🗓️ Date: Feb 4 @ 6pm 📍 Imperial College London
001
Reposted by Yves-Alexandre de Montjoye
Scoiattolo @scarnecchia.net · 13/12/2024
Can confirm: there is a reason we have to mask small cell counts, especially around rare diagnoses, even when using aggregated data. @yvesalexandre.bsky.social’s entire body of work is instructive here
scholar.google.com
Yves-Alexandre de Montjoye
‪Associate Professor at Imperial College London‬ - ‪‪Cited by 9,182‬‬ - ‪Privacy‬ - ‪Machine learning‬ - ‪AI Safety‬ - ‪Memorization‬ - ‪Automated attacks‬
0306