Sign in

Marton Szep

@martonszep.bsky.social
5 followers 4 following 4 posts

I am a PhD candidate at TUM. My doctoral research in Medical NLP develops reliable LLM-based methods for clinical decision support and workflow optimization under the constraints of domain specificity and limited, sensitive clinical data.

PostsRepliesMedia
Reposted by Marton Szep
Differential Privacy Papers @dppapers.bsky.social · 27/01/2026
Unintended Memorization of Sensitive Information in Fine-Tuned Language Models Marton Szep, Jorge Marin Ruiz, Georgios Kaissis, Paulina Seidl, Rüdiger von Eisenhart-Rothe, Florian Hinterwimmer, Daniel Rueckert arxiv.org/abs/2601.17480
Unintended Memorization of Sensitive Information in Fine-Tuned Language Models

Marton Szep, Jorge Marin Ruiz, Georgios Kaissis, Paulina Seidl, Rüdiger von Eisenhart-Rothe, Florian Hinterwimmer, Daniel Rueckert

http://arxiv.org/abs/2601.17480

Fine-tuning Large Language Models (LLMs) on sensitive datasets carries a substantial risk of unintended memorization and leakage of Personally Identifiable Information (PII), which can violate privacy regulations and compromise individual safety. In this work, we systematically investigate a critical and underexplored vulnerability: the exposure of PII that appears only in model inputs, not in training targets. Using both synthetic and real-world datasets, we design controlled extraction probes to quantify unintended PII memorization and study how factors such as language, PII frequency, task type, and model size influence memorization behavior. We further benchmark four privacy-preserving approaches including differential privacy, machine unlearning, regularization, and preference alignment, evaluating their trade-offs between privacy and task performance. Our results show that post-training methods generally provide more consistent privacy-utility trade-offs, while differential privacy achieves strong reduction in leakage in specific settings, although it can introduce training instability. These findings highlight the persistent challenge of memorization in fine-tuned LLMs and emphasize the need for robust, scalable privacy-preserving techniques.
011
Marton Szep @martonszep.bsky.social · 10/03/2026
Thrilled to present our paper "Unintended Memorization of Sensitive Information in Fine-Tuned Language Models" at #EACL2026 in Rabat! 🇲🇦 w/ J. Marin Ruiz, G. Kaissis, P. Seidl, R. v. Eisenhart-Rothe, F. Hinterwimmer & @danielrueckert.bsky.social. Read here: arxiv.org/abs/2601.174...
A promotional graphic for an oral presentation at the EACL 2026 conference in Morocco. The background features a sunny, historic Moroccan stone fortress gate with palm trees, a clear blue sky, and decorative geometric tile patterns in the corners. Text in the top left indicates the event is at Palais Des Congres, Rabat, from March 24-29, 2026. A banner across the middle displays the presentation title: "Unintended Memorization of Sensitive Information in Fine-Tuned Language Models." Below the title is a flowchart diagram illustrating how Large Language Models (LLMs) trained on sensitive medical text can inadvertently memorize Personally Identifiable Information (PII), and how a "True-Prefix Attack" can extract a patient's name even when fine-tuned for downstream tasks that do not contain PII. Text at the very bottom reads, "Oral Presentation: March 27 | 11:00 AM | Salle La Palmeraie."
252