Sign in

Luc Rocher

@rocher.lc
1.9K followers 350 following 207 posts

associate professor at Oxford · UKRI future leaders fellow · i study how data and algorithms shape societies · AI fairness, accountability and transparency · algorithm auditing · photographer, keen 🚴🏻 · they/them · rocher.lc (views my own)

PostsRepliesMedia
Luc Rocher @rocher.lc · 08/06/2026
Don’t think this is a particularly strong study to translate to Ox though. It really depends on specific socioeconomics in London in and out the boundary line (which I’m sure wasn’t randomly drawn). See their main figure.
Figure from study showing nb of registered cars going up over time in and outside zone
120
Luc Rocher @rocher.lc · 29/04/2026
💬 Getting the ick with chatbots? Our research in Nature shows that models trained to adopt a warm and friendly tone are more likely to make mistakes when answering your questions. Small changes in tone can undermine accuracy across architectures, particularly when users express vulnerability.
Banner image with screenshot of paper on the left and first figure on the right, showing two conversations: one with a neutral model, one with a warm model.
1218
Luc Rocher @rocher.lc · 24/04/2026
UK Biobank data was on sale on Alibaba. It was also exposed 197 times in a year, according to my investigations. I talked to The Telegraph about this case: “it’s quite easy still, at the moment, to find UK Biobank data that is there, uploaded, often by mistake, by researchers around the world”.
Screenshot of headline from The Telegraph:

Biobank data leaked 198 times in past year

Researchers had to be ordered to take down volunteers’ medical information they had uploaded to internet 

Confidential medical data held by UK Biobank has been leaked online at least 198 times in the past year.

The Biobank has issued almost 200 legal threats to researchers, urging them to take down the unlawful publication of the health data of thousands of British people, according to experts tracking the breaches.
110
Luc Rocher @rocher.lc · 23/04/2026
Found out today that @nytimes.com is using an AI chatbot which attempts to pass as a real human in their customer service chat…
010
Luc Rocher @rocher.lc · 14/04/2026
New BMJ editorial w/ @jessrmorley.bsky.social ! A new privacy scandal has engulfed UK Biobank, one of the world’s most important biomedical research resources. Researchers have repeatedly & accidentally uploaded datasets to GitHub, forcing Biobank to issue 103 legal takedown notices as of today.
161
Luc Rocher @rocher.lc · 24/02/2026
hello from IASEAI, lots of interest for making safe and ethical AI here 👀
photo of poster at the conference
020
Luc Rocher @rocher.lc · 09/02/2026
⚠️ Despite all the hype, chatbots still make terrible doctors. Out today is the largest user study of language models for medical self-diagnosis. We found that chatbots provide inaccurate and inconsistent answers, and that people are better off using online searches or their own judgment.
Banner image with screenshot of scientific article from nature Medicine, as well as two panels from the study method and results
7357163
Luc Rocher @rocher.lc · 09/02/2026
Roxana Radu and I wrote a short letter about the need to uphold the sanctity of citations in academia, and not giving in to AI. The craft and art of attributing and situating knowledge are in danger, and we believe this would first harm under-credited, often erased work.
Screenshot from article that starts as follow:

Meticulous citation is a marker of well-researched, serious scholarship. Citations do a lot more than attributing credit; they situate claims within the context of existing research and enable scrutiny. When authors cite carelessly, for example by referencing famous figures and articles while overlooking original sources, they make two important errors. First, they credit ideas to the wrong person and, second, they reveal a limited understanding of the relevant scholarship. Misattribution disproportionately harms underrepresented voices, whose work has been shown to be consistently more innovative than that of established researchers1. Research led by women tends to be less cited than comparable research led by men. Similarly, research from underrepresented groups, as well as from amateurs and beginners, tends to lead to more breakthroughs and innovation, yet remains less cited than follow-up work from established researchers.
1175
Luc Rocher @rocher.lc · 15/11/2025
📸 New research just out revealing disturbing problems with AI image captioning. AI models can now analyse both texts and images, and are trained on vast collections of human-created content spanning centuries. Some claims these models could soon help human historians interpret and explain the past.
2126
Luc Rocher @rocher.lc · 04/11/2025
Is there a scientific crisis in AI evaluations‽ We did the hard work of reviewing all recent AI benchmarks in top AI conferences. Amongst the 445 reviewed benchmarks, few offer rigorous evaluations of AI capabilities, making it difficult to actually track advances in the field and compare models.
Screenshot of paper with graphs showing research findings
1112
Luc Rocher @rocher.lc · 18/08/2025
Facial recognition just passed 99.95% accuracy in the latest evaluations by the US National Institute of Standards and Technology. But did you know that these numbers come from pristine lab conditions with small datasets, good lightning, and clear photos?
152
Luc Rocher @rocher.lc · 01/08/2025
🧵New research! People now use AI chatbots for therapy, friendship, and romance. These days, both specialised and general-purpose chatbots are built with a friendly, empathetic tone to best comfort users. We find this new trend is not just cosmetic but seriously harms user safety.
Abstract and figure from preprint linked below
1133
Luc Rocher @rocher.lc · 30/06/2025
just received a promotion to associate professor, so i took this opportunity to redo my website (rocher.lc, feedback welcomed)! it looks a bit more boring that the previous version, but easier to read. and thank you everyone who supported me through the years, wouldn't have made it here with you! 🌱
180
Luc Rocher @rocher.lc · 29/06/2025
@sofiahafner.bsky.social gave an excellent talk about our work, which shows that as language models get larger and bigger, they also learn more stereotypical, binary associations between gender and sex. More here: arxiv.org/abs/2505.14080
030
Luc Rocher @rocher.lc · 29/06/2025
Had a lovely time at FAccT in Athens earlier this week. Attached photos of me, beans, cat hotels, lime&basil ice cream.
140
Luc Rocher @rocher.lc · 14/02/2025
How is your Friday going? This is me trying to figure out how to get a new French passport after the UK Home Office lost it somewhere in Swindon…
030
Luc Rocher @rocher.lc · 21/01/2025
Not often do we see an official FRS, Fellow of the Royal Society, performing a heil hitler salute live in front of such a large audience. Surely the Royal Society is already investigating and will take a prompt decision in the coming days.
2188
Luc Rocher @rocher.lc · 13/01/2025
Mixed feelings about this new AI Opportunities Action Plan. Surprised to see how weak the data strategy is. Some points just don't make sense, e.g. pushing for a single form of privacy-preserving access to sensitive data (synthetic data) which has been shown to provide very little usability.
150
Luc Rocher @rocher.lc · 13/01/2025
Curious to see how they plan to anonymise our NHS data for training ML models. I personally wouldn’t know how to start doing that www.theguardian.com/politics/202...
The government plan features a potentially controversial scheme to unlock public data to help fuel the growth of AI businesses. This includes anonymised NHS data, which will be available for “researchers and innovators” to train their AI models. The government says there would be “strong privacy-preserving safeguards” and the data would never be owned by private companies.
1189
Luc Rocher @rocher.lc · 09/01/2025
While they are likely even more accurate models out there, this scaling law is already much better than previously-used heuristics using exponential decay or polynomials. Interestingly, it's also far better than relying on entropy (think “33 bits of information enough to identify anyone on earth”).
We perform measurement-based extrapolation of (a) exact, (b) sparse, and (c) robust matching attacks. We report the performance of the PYC-MB method compared to three other functional forms (ENT—Entropy baseline with no tail complexity, orange line; EXP—exponential decay function, green line; POL—polynomial function, red line; see Supplementary Note S3.2). We report the performance when trained on (a) exact matching (ADULT-1, using discrete demographics), (b) sparse matching (APPS-1, using 2 installed Android apps), (c) robust matching (GEO-1, ML-based mobile phone geolocation matching). We measure the empirical correctness up to μ ∈ {1%, 5%, 10%} of the original data and, for each sampling fraction μ, fit the four functional forms. We display the fitted correctness with solid color lines and the training part with a gray background. We display the empirical correctness with black dots. In all examples, the PYC-MB achieves high accuracy with good model specification.
110
Luc Rocher @rocher.lc · 09/01/2025
During my PhD, I taught myself Bayesian statistics and discovered a beautiful field of work on nonparametrics models. That's what we now exploit here to characterise what makes individuals identifiable, resulting in a simple scaling law for the {accuracy/correctness/rank-1 identification rate}.
Visualisation of the experimental setting with a gallery of identifiable (or not) individuals, and the Pitman-Yor Correctness model that we designed
110
Luc Rocher @rocher.lc · 09/01/2025
Entire fields test biometrics and human id. models on small-scale benchmarks, reporting accuracy, AUC, true positive rate, etc. that do not match performance in real-world where id. is much harder. Worse, a model A better than a model B in small-scale isn't necessarily better at scale!
We report the expected correctness in two scenarios for a fixed world population of n = 7.53 billion people. a Effect of the tail complexity parameter γ on the expected correctness. Each line represents the correctness for a fixed entropy h from 10 to 60 bits, with color indicating the entropy h. b Effect of the entropy h on the expected correctness. Critical behaviors arise for exponential tails (γ = 0, top) but not for heavy tails (γ = 0.5, middle; γ = 1, bottom).
110
Luc Rocher @rocher.lc · 09/01/2025
🔴 Finally out, my work in @naturecomms.bsky.social investigates (re-)identification techniques from browser fingerprinting to facial recognition. It's a surprisingly tough problem, with incorrect heuristics and rules of thumbs often used—huge issue for accountability and independent investigations.
Each panel shows the empirical correctness κ (black dots) in four identification scenarios, along with our prediction (solid blue line) fitted on the empirical κ scores. a Identification of mobile phone users from their pseudonymized 1-hop social network (IIG-1, n = 43, 000 phones) by Creţu et al. b Facial recognition using Google FaceNet V8 (FACEREC-2, n = 1M faces) by Kemelmacher-Shlizerman et al. c Authorship attribution in textual data using Deep Learning (TEXT-1, n = 500 authors) by Saedi et al. d Exact matching using simple browser fingerprints (HTTP accept, cookies and JavaScript enabled, timezone, display size, installed fonts, plugins, user agent, video) collected by Panopticlick (WEB-2, n = 5.5M fingerprints).
193
Luc Rocher @rocher.lc · 21/12/2024
Surprised that private actors are even allowed to participate in digital ID efforts. France has had one single app for years, it’s a very simple app, it works fine, no need for dozens of them.
Quote from article about Government sanctioning independent apps
100
Luc Rocher @rocher.lc · 13/12/2024
You might say, well, at least they can hardly identify users, right? Right?
Screenshot of the news release with the following sentence highlighted: “Our Trust and Safety teams can use this bottom-up review approach to identify individual accounts for further review and, if appropriate, take action in accordance with our terms and policies.”
000
Luc Rocher @rocher.lc · 13/12/2024
To comply with their privacy policy, Anthropic built a new privacy-preserving analytics platform that creates anonymized insights from all free and Pro Claude customers' chats. How do they look for any privacy risks? They just ask Claude.
130
Luc Rocher @rocher.lc · 12/12/2024
Adding bureaucratic requirements to a burdened system risks leaving the NHS vulnerable to exploitation by private technology companies whose offers to “assist” with infrastructure development could result in loss of control over valuable public assets. @oiioxford.bsky.social @yaledec.bsky.social
Adding bureaucratic requirements to an already burdened system, while perpetuating ideas about NHS technical limitations and false dichotomies will slow progress. This risks leaving the NHS vulnerable to exploitation by private technology companies whose offers to “assist” with infrastructure development could result in loss of control over valuable public assets. Instead, the NHS needs strategic investment in teams capable of developing privacy preserving platforms at scale with robust security measures that go beyond criminalising re-identification. This technical innovation, supported by proper evaluation frameworks to verify security claims and ensure research integrity, can deliver the “critical national infrastructure” the review supports.
120
Luc Rocher @rocher.lc · 12/12/2024
New editorial in @bmj.com by @jessrmorley.bsky.social and myself! Technical solutions are needed to modernise the NHS health data infrastructure. Investment in privacy preserving platforms with robust security measures is key. 🔗 bmj.com/cgi/content/full/bmj.q2735?ijkey=NzzxQSQlpEvi3Mx
Technical, rather than bureaucratic, solutions are needed

Reforming the NHS by shifting from analogue to digital, from treating sickness to prevention of disease, and from hospital to community care is a priority for the UK government.1 Better use of data will be central to achieving these shifts—revealing who is likely to become unwell, enabling predictive modelling, and simulating the effects of changing the location of care. The November 2024 publication of the Sudlow review of the UK’s health data systems2 is therefore timely. Commissioned by the chief medical officer for England, the review makes recommendations for overcoming barriers to linking and sharing data by streamlining control; standardising mechanisms, governance policies, and public engagement activities for data access; and broadening access to imaging and free text data.

Infrastructure needs to be consistent and better coordinated. Yet recommendations in the Sudlow review for creating “critical national infrastructure” over-rely on bureaucratic solutions (rather than technical ones) and fail to resolve key tensions around privacy, public benefit, data security, and trust.
1156
Luc Rocher @rocher.lc · 11/12/2024
Only 11% believe that AI tools have reduced their workload. All results in French: solidairesfinancespubliques.org/vie-des-serv...
010
Luc Rocher @rocher.lc · 16/11/2024
Yes! The traditionnal version with almond cream plus crème pâtissière. 🥰 Attached my notes on galettes, it has the most fascinating history.
Rooted in an equalitarian tradition of reversal of hierarchies, but also a symbol of fertility and arrival of the spring, the galette holds many secrets. Collection welcomed from North London until the end of the month.
A French traditional pastry, this galette is an old chestnut of January you will find in every respectful bakery of the kingdom. It is made of two puff pastry discs enclosing a frangipane (1/3 crème patissière, 2/3 almond cream). I used to skip the crème patissière, and one should not (otherwise, you end up with a Pithivier, and not a galette). This delicious and healthy conglomerate of wheat, butter, and almonds originates from ancestral Roman traditions *but* also played a key role in the French anti-monarchist and republican saga since the middle age.
The modern galette is made with puff pastry (leafed pastry). It has a controversial history, as many French recipes. Many claim puff pastry was invented by Claude Gelée in 1645. This is likely wrong! Books from 1311 already mention “gasteaux feuillés”; others point to musammana in Moorish Spain as the origin of layered pastry ca 13th century too.
That said, we can also trace back the galette to the roman Saturnalias, where fava beans were hidden inside king cakes, and whoever found the bean would be king of the day. Similar traditions seem to have reached England, where whoever receives the bean-ed slice is crowned Lord of Misrule (King of the Bean) and apparently Queen of the Pea if you find a pea instead of a bean?
Less know is finally the anti-monarchist tradition associated with this galette in France. To eat the galette, one must first « tirer les rois »—the youngest person hides under the table to randomly assigning slices to guests. The monarch (« roi de la fève ») is thus randomly elected. This theme was widely used in popular culture to critique and undermine the monarchy.
120
Luc Rocher @rocher.lc · 16/11/2024
I nominate myself! 👋
1101
Luc Rocher @rocher.lc · 12/11/2024
📝 A few weeks ago, I started a checklist for scientific writing and editing. I cover how to set a narrative; improve paragraph structure/transitions/signposting; write in a simple and precise way; and organise introduction, results, discussion, and methods. Feedback is very welcome! osf.io/4cyek
073
Luc Rocher @rocher.lc · 11/11/2024
Never seen any so far. Does that still happen when hiding what they call "non-sexual nudity"? Might be images flagged as light nudity but w/o content warning.
110
Luc Rocher @rocher.lc · 09/07/2023
Alright 👋 just got in, hello there!
050