Sign in

Michael Hahn

@m-hahn.bsky.social
112 followers 253 following 8 posts

Prof at Saarland. NLP and machine learning. Theory and interpretability of LLMs. www.mhahn.info

PostsRepliesMedia
Reposted by Michael Hahn
micha heilbron @mheilbron.bsky.social · 10/03/2026
📢 PhD position in Developmental Language Modelling (PLZ RT) What can human language acquisition teach us about training language models? Join us as a PhD! mpi.nl/career-education/vacancies/vacancy/fully-funded-4-year-phd-position-developmental-language @carorowland.bsky.social @mpi-nl.bsky.social
12935
Reposted by Michael Hahn
Harvey Fu @harveyfu.bsky.social · 20/06/2025
LLMs excel at finding surprising “needles” in very long documents, but can they detect when information is conspicuously missing? 🫥AbsenceBench🫥 shows that even SoTA LLMs struggle on this task, suggesting that LLMs have trouble perceiving “negative spaces”. Paper: arxiv.org/abs/2506.11440 🧵[1/n]
27416
Reposted by Michael Hahn
Guy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 23/05/2025
New preprint alert! We often prompt ICL tasks using either demonstrations or instructions. How much does the form of the prompt matter to the task representation formed by a language model? Stick around to find out 1/N
1457
Michael Hahn @m-hahn.bsky.social · 05/05/2025
Chain-of-Thought (CoT) reasoning lets LLMs solve complex tasks, but long CoTs are expensive. How short can they be while still working? Our new ICML paper tackles this foundational question.
2122
Reposted by Michael Hahn
Marius Mosbach @mariusmosbach.bsky.social · 09/04/2025
Check out our new paper on unlearning for LLMs 🤖. We show that *not all data are unlearned equally* and argue that future work on LLM unlearning should take properties of the data to be unlearned into account. This work was lead by my intern @a-krishnan.bsky.social 🔗: arxiv.org/abs/2504.05058
Diagram illustrating a hypothesis about knowledge unlearning in language models. The left side shows a training corpus with varying frequencies of facts, such as 'Montreal is a city in Quebec' (high frequency) and 'Atlantis is a city in the ocean' (lower frequency). The center shows a language model being trained on this data, then undergoing unlearning. The right side demonstrates the 'Forget Quality' results, where the model more effectively unlearns the less frequent fact ('Atlantis is in Greece') while retaining the more frequent knowledge. Labels A, B, and C mark key points in the hypothesis: A (frequency variations in training data), B (influence of frequency), and C (unlearning effectiveness).
1325
Reposted by Michael Hahn
Surya Ganguli @suryaganguli.bsky.social · 31/12/2024
Our new paper! "Analytic theory of creativity in convolutional diffusion models" lead expertly by @masonkamb.bsky.social arxiv.org/abs/2412.20292 Our closed-form theory needs no training, is mechanistically interpretable & accurately predicts diffusion model outputs with high median r^2~0.9
513431
Reposted by Michael Hahn
Preetum Nakkiran @preetumnakkiran.bsky.social · 01/01/2025
Just read this, neat paper! I really enjoyed Figure 3 illustrating the basic idea: Suppose you train a diffusion model where the denoiser is restricted to be "local" (each pixel i only depends on its 3x3 neighborhood N(i)). The optimal local denoiser for pixel i is E[ x_0[i] | x_t[ N(i) ] ]...cont
1392