Sign in

Valentin Hofmann

@valentinhofmann.bsky.social
2.4K followers 169 following 39 posts

Assistant Professor @cislmu.bsky.social @lmu.de

PostsRepliesMedia
Reposted by Valentin Hofmann
Carolin Holtermann @carolin-holtermann.bsky.social · 24/03/2026
🚨 New paper alert: A new generation of LLMs can now process speech natively. This could expand access for millions excluded by text interfaces, but our research shows a cost: demographic cues in speaker voice can trigger stereotypical model responses. 🎙️⚖️ Paper: arxiv.org/abs/2603.22260
3175
Valentin Hofmann @valentinhofmann.bsky.social · 03/03/2026
📢 Life update 📢 After a wonderful time at @ai2.bsky.social, I've joined @cislmu.bsky.social at @lmu.de as a tenure-track assistant professor in NLP. Thrilled to be back in Europe and to start a lab in Munich's flourishing AI ecosystem! 🎉
2291
Reposted by Valentin Hofmann
Manuel Tonneau @manueltonneau.bsky.social · 27/01/2026
Demographic cues (eg, names, dialect) are widely used to study how LLM behavior may change depending on user demographics. Such cues are often assumed interchangeable. 🚨 We show they are not: different cues yield different model behavior for the same group and different conclusions on LLM bias. 🧵👇
1189
Reposted by Valentin Hofmann
Ai2 @ai2.bsky.social · 15/12/2025
Introducing Bolmo, a new family of byte-level language models built by "byteifying" our open Olmo 3—and to our knowledge, the first fully open byte-level LM to match or surpass SOTA subword models across a wide range of tasks. 🧵
17315
Valentin Hofmann @valentinhofmann.bsky.social · 31/10/2025
Excited to see our #COLM2025 paper on fluid benchmarking highlighted by @eval-eval.bsky.social! They are worth a follow if you are into LLM eval research. 🔬
020
Reposted by Valentin Hofmann
Paul Röttger @paul-rottger.bsky.social · 29/10/2025
There’s plenty of evidence for political bias in LLMs, but very few evals reflect realistic LLM use cases — which is where bias actually matters. IssueBench, our attempt to fix this, is accepted at TACL, and I will be at #EMNLP2025 next week to talk about it! New results 🧵
13111
Valentin Hofmann @valentinhofmann.bsky.social · 14/10/2025
Check out this #EMNLP2025 paper led by @minhducbui.bsky.social and @carolin-holtermann.bsky.social showing dialect prejudice remains a major issue in current LLMs. Example: GPT-5 associates German dialect speakers with being uneducated and steers them toward stereotyped jobs (e.g., farmworkers). 👇
070
Valentin Hofmann @valentinhofmann.bsky.social · 19/09/2025
Thanks, Jordan! Your ACL 2021 paper was a huge source of inspiration for us!
010
Valentin Hofmann @valentinhofmann.bsky.social · 19/09/2025
We did not specifically analyze novel models as your paper did. While I am optimistic that Fluid Benchmarking improves over static IRT-based methods in this regime as well, there are definitely limitations, which we discuss in the paragraph below. Would be exciting to run more experiments on this!
000
Valentin Hofmann @valentinhofmann.bsky.social · 19/09/2025
In our experiments, we find that this dynamic approach consistently outperforms static IRT-based methods. The improvements are especially pronounced in terms of variance, which poses a major challenge for static IRT-based methods. We discuss this in more detail in the paragraph below.
100
Valentin Hofmann @valentinhofmann.bsky.social · 19/09/2025
Great question! The key difference is that we use IRT to dynamically adapt the subset of items to a model's capability, rather than to determine a static, "globally optimal" subset of items as in prior work. With Fluid Benchmarking, each model is evaluated on a different subset of items.
100
Reposted by Valentin Hofmann
Kyle Lo @ COLM2026 @kylelo.bsky.social · 17/09/2025
LM benchmark design requires 3 decisions, how to: 🐟 select test cases 🐠 score LM on each test 🦈 aggregate scores to estimate perf fluid benchmarking is simple: 🍣 find max informative test cases 🍥 estimate 'ability', not simple avg perf why care? turn ur grey noisy benchmarks to red ones!
052
Valentin Hofmann @valentinhofmann.bsky.social · 16/09/2025
Last but not least, a huge shoutout to my incredible coauthors @davidheineman.com, @ianmagnusson.bsky.social, @kylelo.bsky.social, @jessedodge.bsky.social, @maartensap.bsky.social, Pang Wei Koh, Chun Wang, @hanna-nlp.bsky.social, and @nlpnoah.bsky.social! 🤗
030
Valentin Hofmann @valentinhofmann.bsky.social · 16/09/2025
For details, check out our paper, blog, code, and data: 📄 arxiv.org/abs/2509.11106 ✍️ allenai.org/blog/fluid-b... 💻 github.com/allenai/flui... 📊 huggingface.co/datasets/all... Looking forward to chatting more at #COLM2025! 👋
120
Valentin Hofmann @valentinhofmann.bsky.social · 16/09/2025
Overall, our work shows that LLM evaluations can be substantially improved by moving beyond the until-now universal practice of static benchmarking, which assumes a globally optimal set of evaluation questions for all models.
120
Valentin Hofmann @valentinhofmann.bsky.social · 16/09/2025
These (and more) advantages are achieved while at the same time reducing evaluation cost. Example: on MMLU, Fluid Benchmarking results in lower step-to-step variance and higher validity than standard methods while using 50 times fewer questions. ⚡
110
Valentin Hofmann @valentinhofmann.bsky.social · 16/09/2025
Fluid Benchmarking substantially reduces step-to-step variance during pretraining. It also increases validity: results generalize better to other benchmarks targeting the same capability. One reason: it automatically avoids mislabeled questions, cutting label errors by 99%! 🤯
110
Valentin Hofmann @valentinhofmann.bsky.social · 16/09/2025
In our experiments, we apply Fluid Benchmarking to evaluation during pretraining, a setting where capabilities evolve rapidly. We find that Fluid Benchmarking dynamically adapts to these changes, administering easier questions early in training and more difficult ones later.
110
Valentin Hofmann @valentinhofmann.bsky.social · 16/09/2025
Fluid Benchmarking repeats this loop until the number of administered questions reaches the allotted budget. Adaptive question selection means that LLMs face different sets of questions, but ability estimation aligns results in a common space.
110
Valentin Hofmann @valentinhofmann.bsky.social · 16/09/2025
In Fluid Benchmarking, we start with an initial ability estimate from one question. To select the next question, we use Fisher information. Essentially: a question close in difficulty (b) to the ability estimate (θ) and with high discrimination (a). Then we update the estimate.
110
Valentin Hofmann @valentinhofmann.bsky.social · 16/09/2025
In addition, IRT models each LLM's ability, which can be estimated from its responses to questions with known difficulty and discrimination. The IRT ability estimate can be used to summarize performance like accuracy, and it accounts for question characteristics.
100
Valentin Hofmann @valentinhofmann.bsky.social · 16/09/2025
To get a question's difficulty, we use item response theory (IRT): we analyze responses of hundreds of LLMs to see how often a question is answered correctly. IRT also measures the discrimination of a question, meaning how reliably it separates stronger from weaker LLMs.
110
Valentin Hofmann @valentinhofmann.bsky.social · 16/09/2025
Test theory says: questions are most informative when matched to a test taker's ability. For LLMs, that means evaluating weaker models on easier questions and stronger models on harder ones. But how do we know a question's difficulty, or an LLM's ability, before evaluation? 🤔
110
Valentin Hofmann @valentinhofmann.bsky.social · 16/09/2025
📢 New #COLM2025 paper 📢 Standard benchmarks give every LLM the same questions. This is like testing 5th graders and college seniors with *one* exam! 🥴 Meet Fluid Benchmarking, a capability-adaptive eval method delivering lower variance, higher validity, and reduced cost. 🧵
34110
Reposted by Valentin Hofmann
Dallas Card @dallascard.bsky.social · 29/07/2025
I am delighted to share our new #PNAS paper, with @grvkamath.bsky.social @msonderegger.bsky.social and @sivareddyg.bsky.social, on whether age matters for the adoption of new meanings. That is, as words change meaning, does the rate of adoption vary across generations? www.pnas.org/doi/epdf/10....
35013
Valentin Hofmann @valentinhofmann.bsky.social · 16/07/2025
Attending #ICML2025? Don't miss this TokShop panel, which will explore: 🔮 The Future of Tokenization 🔮 Featuring a stellar lineup of panelists - mark your calendar! ✨
040
Valentin Hofmann @valentinhofmann.bsky.social · 10/06/2025
LLMs can appear unbiased on the surface but still perpetuate racist views in subtle ways. What causes this discrepancy? 🔍 In our upcoming #ACL2025 paper, we find a pattern akin to racial colorblindness: LLMs suppress race in ambiguous contexts, leading to biased outcomes.
060
Reposted by Valentin Hofmann
Tokenization Workshop (TokShop) @COLM2026 @tokshop.bsky.social · 30/05/2025
📣 We extend the submission deadline by 24 hours to avoid conflict with ACL camera-ready deadline. 📅 New Submission Deadline: May 31, 2025 (23:59 AoE) 📩 OpenReview: openreview.net/group?id=ICM...
openreview.net
ICML 2025 Workshop TokShop
Welcome to the OpenReview homepage for ICML 2025 Workshop TokShop
011
Valentin Hofmann @valentinhofmann.bsky.social · 29/05/2025
Huge congrats, Adam!!! 🎉
010
Reposted by Valentin Hofmann
Tokenization Workshop (TokShop) @COLM2026 @tokshop.bsky.social · 28/05/2025
Got a good tokenization paper under review at COLM, but the scores were a letdown? 😬 Why bother with rebuttal when the perfect venue is right around the corner! Submit your paper to the #ICML2025 Tokenization Workshop (TokShop) by May 30! 🚀
0104
Reposted by Valentin Hofmann
Tokenization Workshop (TokShop) @COLM2026 @tokshop.bsky.social · 26/05/2025
Beyond text: Modern AI tokenizes images too! Vision models split photos into patches, treating each 16x16 pixel square as a "token." 🖼️➡️🔤 #VisualTokenization Interested in tokenization? Join our workshop tokenization-workshop.github.io The submission deadline is already May 30!
tokenization-workshop.github.io
042
Valentin Hofmann @valentinhofmann.bsky.social · 10/05/2025
Yes, exactly! And we make sure that the words do not appear in the language model's pretraining data.
010
Reposted by Valentin Hofmann
Ai2 @ai2.bsky.social · 09/05/2025
Do LLMs learn language via rules or analogies? This could be a surprise to many – models rely heavily on stored examples and draw analogies when dealing with unfamiliar words, much as humans do. Check out this new study led by @valentinhofmann.bsky.social to learn how they made the discovery 💡
1225
Valentin Hofmann @valentinhofmann.bsky.social · 09/05/2025
Check out the published version of the paper here: www.pnas.org/doi/10.1073/...
pnas.org
PNAS
Proceedings of the National Academy of Sciences (PNAS), a peer reviewed journal of the National Academy of Sciences (NAS) - an authoritative source of high-impact, original research that broadly spans...
030
Valentin Hofmann @valentinhofmann.bsky.social · 09/05/2025
Thrilled to share that this is out in @pnas.org today! 🎉 We show that linguistic generalization in language models can be due to underlying analogical mechanisms. Shoutout to my amazing co-authors @weissweiler.bsky.social, @davidrmortensen.bsky.social, Hinrich Schütze, and Janet Pierrehumbert!
1356
Reposted by Valentin Hofmann
Tokenization Workshop (TokShop) @COLM2026 @tokshop.bsky.social · 04/05/2025
Got a tokenization paper that just didn't make the cut for ICML? Submit it to the Tokenization Workshop TokShop at #ICML2025 -- we'd love to see it there! tokenization-workshop.github.io
tokenization-workshop.github.io
Tokenization Workshop @ ICML 2025
076
Reposted by Valentin Hofmann
Ben Waber @bwaber.bsky.social · 25/04/2025
Next was a fantastic talk by @valentinhofmann.bsky.social on probing covert racism in LLMs at the @ltiatcmu.bsky.social. One can imagine where this goes. Highly recommend www.youtube.com/watch?v=_1Ej... (5/11)
youtube.com
4.18.25 LTI Colloquium Valentin Hofmann
YouTube video by Language Technologies Institute at Carnegie Mellon (LTI at CMU)
152
Valentin Hofmann @valentinhofmann.bsky.social · 15/04/2025
Delighted there will finally be a workshop devoted to tokenization - a critical topic for LLMs and beyond! 🎉 Join us for the inaugural edition of TokShop at #ICML2025 @icmlconf.bsky.social in Vancouver this summer! 🤗
0265
Reposted by Valentin Hofmann
Benjamin Minixhofer @bminixhofer.bsky.social · 02/04/2025
We created Approximate Likelihood Matching, a principled (and very effective) method for *cross-tokenizer distillation*! With ALM, you can create ensembles of models from different families, convert existing subword-level models to byte-level and a bunch more🧵
Image illustrating that ALM can enable Ensembling, Transfer to Bytes, and general Cross-Tokenizer Distillation.
12514
Valentin Hofmann @valentinhofmann.bsky.social · 21/03/2025
Humans store thousands of multi-word expressions like "of course" in their mental lexicon, but current tokenizers don't support multi-word tokens. Enter SuperBPE, a tokenizer that lifts this restriction and brings substantial gains in efficiency and performance! 🚀 Details 👇
151
Reposted by Valentin Hofmann
Julia Mendelsohn @jmendelsohn2.bsky.social · 20/02/2025
New preprint! Metaphors shape how people understand politics, but measuring them (& their real-world effects) is hard. We develop a new method to measure metaphor & use it to study dehumanizing metaphor in 400K immigration tweets Link: bit.ly/4i3PGm3 #NLP #NLProc #polisky #polcom #compsocialsci 🐦🐦
Screenshot of top half of first page of paper. The paper is titled: "When People are Floods: Analyzing Dehumanizing Metaphors in Immigration Discourse with Large Language Models". The authors are Julia Mendelsohn (University of Chicago) and Ceren Budak (University of Michigan). The top right corner contains a visual showing the sentence "They want immigrants to pour into and infest this country". The caption says: Figure 1: Dehumanizing sentence likening immigrants to the source domain concepts of Water and Vermin via the words "pour" and "infest". 

The abstract text on the left reads: Metaphor, discussing one concept in terms of another, is abundant in politics and can shape how people understand important issues. We develop a computational approach to measure metaphorical language, focusing on immigration discourse on social media. Grounded in qualitative social science research, we identify seven concepts evoked in immigration discourse (e.g. "water" or "vermin"). We propose and evaluate a novel technique that leverages both word-level and document-level signals to measure metaphor with respect to these concepts. We then study the relationship between metaphor, political ideology, and user engagement in 400K US tweets about immigration. While conservatives tend to use dehumanizing metaphors more than liberals, this effect varies widely across concepts. Moreover, creature-related metaphor is associated with more retweets, especially for liberal authors. Our work highlights the potential for computational methods to complement qualitative approaches in understanding subtle and implicit language in political discourse.
618264
Reposted by Valentin Hofmann
Leonie Weissweiler @weissweiler.bsky.social · 20/02/2025
✨New paper✨ Linguistic evaluations of LLMs often implicitly assume that language is generated by symbolic rules. In a new position paper, @adelegoldberg.bsky.social, @kmahowald.bsky.social and I argue that languages are not Lego sets, and evaluations should reflect this! arxiv.org/pdf/2502.13195
16819
Valentin Hofmann @valentinhofmann.bsky.social · 13/02/2025
Excited to share IssueBench, the most extensive benchmark for LLM political bias! 📊 We find surprising consistency across models, with notable differences in Qwen on China-related issues. All examined LLMs also show strong alignment with Democrat voters. More details below! 👇
0101
Valentin Hofmann @valentinhofmann.bsky.social · 31/01/2025
Great to see the International AI Safety Report highlight research on dialect prejudice, including our work on covert racism in LLMs! www.nature.com/articles/s41...
161
Reposted by Valentin Hofmann
Paul Röttger @paul-rottger.bsky.social · 21/01/2025
Today, we are releasing MSTS, a new Multimodal Safety Test Suite for vision-language models! MSTS is exciting because it tests for safety risks *created by multimodality*. Each prompt consists of a text + image that *only in combination* reveal their full unsafe meaning. 🧵
23011
Valentin Hofmann @valentinhofmann.bsky.social · 05/12/2024
Check out the paper for many more details and analyses: arxiv.org/abs/2411.07990 This is the last paper from my PhD, and it was a real pleasure to work on it with @weissweiler.bsky.social, @davidrmortensen.bsky.social, Hinrich Schütze, and Janet Pierrehumbert!
arxiv.org
Derivational Morphology Reveals Analogical Generalization in Large Language Models
What mechanisms underlie linguistic generalization in large language models (LLMs)? This question has attracted considerable attention, with most studies analyzing the extent to which the language ski...
181
Valentin Hofmann @valentinhofmann.bsky.social · 05/12/2024
We observe a frequency effect across all adjective classes that gradually gets stronger the more variable an adjective class is. This is again exactly in line with analogical models, where rule-like behavior is the end of a gradient characterized by varying levels of regularity.
120
Valentin Hofmann @valentinhofmann.bsky.social · 05/12/2024
To test this hypothesis, we analyze how frequency affects GPT-J's confidence when only one of the two derivatives is attested in the Pile. Crucially, if GPT-J handled regular adjective classes via rules, then how often it has seen a derivative should not impact its confidence.
120
Valentin Hofmann @valentinhofmann.bsky.social · 05/12/2024
These results suggest that LLMs generalize linguistic phenomena with a high degree of variability via analogical processes. However, there remains the possibility that LLMs effectively use analogies in cases of variation and apply rules for regular adjective classes.
130
Valentin Hofmann @valentinhofmann.bsky.social · 05/12/2024
As expected, rule-based and analogical models make the same predictions for regular adjective classes (-able, -ish) and thus explain GPT-J's behavior equally well. However, for variable adjective classes (-ive, -ous), the analogical model results in a significantly better match.
120