Reposted by Valentin HofmannCarolin Holtermann @carolin-holtermann.bsky.social · 24/03/2026🚨 New paper alert: A new generation of LLMs can now process speech natively. This could expand access for millions excluded by text interfaces, but our research shows a cost: demographic cues in speaker voice can trigger stereotypical model responses. 🎙️⚖️ Paper: arxiv.org/abs/2603.22260 3175
Valentin Hofmann @valentinhofmann.bsky.social · 03/03/2026📢 Life update 📢 After a wonderful time at @ai2.bsky.social, I've joined @cislmu.bsky.social at @lmu.de as a tenure-track assistant professor in NLP. Thrilled to be back in Europe and to start a lab in Munich's flourishing AI ecosystem! 🎉 2291
Reposted by Valentin HofmannManuel Tonneau @manueltonneau.bsky.social · 27/01/2026Demographic cues (eg, names, dialect) are widely used to study how LLM behavior may change depending on user demographics. Such cues are often assumed interchangeable. 🚨 We show they are not: different cues yield different model behavior for the same group and different conclusions on LLM bias. 🧵👇 1189
Reposted by Valentin HofmannAi2 @ai2.bsky.social · 15/12/2025Introducing Bolmo, a new family of byte-level language models built by "byteifying" our open Olmo 3—and to our knowledge, the first fully open byte-level LM to match or surpass SOTA subword models across a wide range of tasks. 🧵 17315
Valentin Hofmann @valentinhofmann.bsky.social · 31/10/2025Excited to see our #COLM2025 paper on fluid benchmarking highlighted by @eval-eval.bsky.social! They are worth a follow if you are into LLM eval research. 🔬 020
Reposted by Valentin HofmannPaul Röttger @paul-rottger.bsky.social · 29/10/2025There’s plenty of evidence for political bias in LLMs, but very few evals reflect realistic LLM use cases — which is where bias actually matters. IssueBench, our attempt to fix this, is accepted at TACL, and I will be at #EMNLP2025 next week to talk about it! New results 🧵 13111
Valentin Hofmann @valentinhofmann.bsky.social · 14/10/2025Check out this #EMNLP2025 paper led by @minhducbui.bsky.social and @carolin-holtermann.bsky.social showing dialect prejudice remains a major issue in current LLMs. Example: GPT-5 associates German dialect speakers with being uneducated and steers them toward stereotyped jobs (e.g., farmworkers). 👇 070
Reposted by Valentin HofmannKyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 17/09/2025LM benchmark design requires 3 decisions, how to: 🐟 select test cases 🐠 score LM on each test 🦈 aggregate scores to estimate perf fluid benchmarking is simple: 🍣 find max informative test cases 🍥 estimate 'ability', not simple avg perf why care? turn ur grey noisy benchmarks to red ones! 052
Valentin Hofmann @valentinhofmann.bsky.social · 16/09/2025📢 New #COLM2025 paper 📢 Standard benchmarks give every LLM the same questions. This is like testing 5th graders and college seniors with *one* exam! 🥴 Meet Fluid Benchmarking, a capability-adaptive eval method delivering lower variance, higher validity, and reduced cost. 🧵 34110
Reposted by Valentin HofmannDallas Card @dallascard.bsky.social · 29/07/2025I am delighted to share our new #PNAS paper, with @grvkamath.bsky.social @msonderegger.bsky.social and @sivareddyg.bsky.social, on whether age matters for the adoption of new meanings. That is, as words change meaning, does the rate of adoption vary across generations? www.pnas.org/doi/epdf/10.... 35013
Valentin Hofmann @valentinhofmann.bsky.social · 16/07/2025Attending #ICML2025? Don't miss this TokShop panel, which will explore: 🔮 The Future of Tokenization 🔮 Featuring a stellar lineup of panelists - mark your calendar! ✨ 040
Valentin Hofmann @valentinhofmann.bsky.social · 10/06/2025LLMs can appear unbiased on the surface but still perpetuate racist views in subtle ways. What causes this discrepancy? 🔍 In our upcoming #ACL2025 paper, we find a pattern akin to racial colorblindness: LLMs suppress race in ambiguous contexts, leading to biased outcomes. 060
Reposted by Valentin HofmannTokenization Workshop (TokShop) @COLM2026 @tokshop.bsky.social · 30/05/2025📣 We extend the submission deadline by 24 hours to avoid conflict with ACL camera-ready deadline. 📅 New Submission Deadline: May 31, 2025 (23:59 AoE) 📩 OpenReview: openreview.net/group?id=ICM...openreview.netICML 2025 Workshop TokShopWelcome to the OpenReview homepage for ICML 2025 Workshop TokShop 011
Reposted by Valentin HofmannTokenization Workshop (TokShop) @COLM2026 @tokshop.bsky.social · 28/05/2025Got a good tokenization paper under review at COLM, but the scores were a letdown? 😬 Why bother with rebuttal when the perfect venue is right around the corner! Submit your paper to the #ICML2025 Tokenization Workshop (TokShop) by May 30! 🚀 0104
Reposted by Valentin HofmannTokenization Workshop (TokShop) @COLM2026 @tokshop.bsky.social · 26/05/2025Beyond text: Modern AI tokenizes images too! Vision models split photos into patches, treating each 16x16 pixel square as a "token." 🖼️➡️🔤 #VisualTokenization Interested in tokenization? Join our workshop tokenization-workshop.github.io The submission deadline is already May 30!tokenization-workshop.github.io 042
Reposted by Valentin HofmannAi2 @ai2.bsky.social · 09/05/2025Do LLMs learn language via rules or analogies? This could be a surprise to many – models rely heavily on stored examples and draw analogies when dealing with unfamiliar words, much as humans do. Check out this new study led by @valentinhofmann.bsky.social to learn how they made the discovery 💡 1225
Valentin Hofmann @valentinhofmann.bsky.social · 09/05/2025Thrilled to share that this is out in @pnas.org today! 🎉 We show that linguistic generalization in language models can be due to underlying analogical mechanisms. Shoutout to my amazing co-authors @weissweiler.bsky.social, @davidrmortensen.bsky.social, Hinrich Schütze, and Janet Pierrehumbert! 1356
Reposted by Valentin HofmannTokenization Workshop (TokShop) @COLM2026 @tokshop.bsky.social · 04/05/2025Got a tokenization paper that just didn't make the cut for ICML? Submit it to the Tokenization Workshop TokShop at #ICML2025 -- we'd love to see it there! tokenization-workshop.github.iotokenization-workshop.github.ioTokenization Workshop @ ICML 2025 076
Reposted by Valentin HofmannBen Waber @bwaber.bsky.social · 25/04/2025Next was a fantastic talk by @valentinhofmann.bsky.social on probing covert racism in LLMs at the @ltiatcmu.bsky.social. One can imagine where this goes. Highly recommend www.youtube.com/watch?v=_1Ej... (5/11)youtube.com4.18.25 LTI Colloquium Valentin HofmannYouTube video by Language Technologies Institute at Carnegie Mellon (LTI at CMU) 152
Valentin Hofmann @valentinhofmann.bsky.social · 15/04/2025Delighted there will finally be a workshop devoted to tokenization - a critical topic for LLMs and beyond! 🎉 Join us for the inaugural edition of TokShop at #ICML2025 @icmlconf.bsky.social in Vancouver this summer! 🤗 0265
Reposted by Valentin HofmannBenjamin Minixhofer @bminixhofer.bsky.social · 02/04/2025We created Approximate Likelihood Matching, a principled (and very effective) method for *cross-tokenizer distillation*! With ALM, you can create ensembles of models from different families, convert existing subword-level models to byte-level and a bunch more🧵 12514
Valentin Hofmann @valentinhofmann.bsky.social · 21/03/2025Humans store thousands of multi-word expressions like "of course" in their mental lexicon, but current tokenizers don't support multi-word tokens. Enter SuperBPE, a tokenizer that lifts this restriction and brings substantial gains in efficiency and performance! 🚀 Details 👇 151
Reposted by Valentin HofmannJulia Mendelsohn @jmendelsohn2.bsky.social · 20/02/2025New preprint! Metaphors shape how people understand politics, but measuring them (& their real-world effects) is hard. We develop a new method to measure metaphor & use it to study dehumanizing metaphor in 400K immigration tweets Link: bit.ly/4i3PGm3 #NLP #NLProc #polisky #polcom #compsocialsci 🐦🐦 618264
Reposted by Valentin HofmannLeonie Weissweiler @weissweiler.bsky.social · 20/02/2025✨New paper✨ Linguistic evaluations of LLMs often implicitly assume that language is generated by symbolic rules. In a new position paper, @adelegoldberg.bsky.social, @kmahowald.bsky.social and I argue that languages are not Lego sets, and evaluations should reflect this! arxiv.org/pdf/2502.13195 16819
Valentin Hofmann @valentinhofmann.bsky.social · 13/02/2025Excited to share IssueBench, the most extensive benchmark for LLM political bias! 📊 We find surprising consistency across models, with notable differences in Qwen on China-related issues. All examined LLMs also show strong alignment with Democrat voters. More details below! 👇 0101
Valentin Hofmann @valentinhofmann.bsky.social · 31/01/2025Great to see the International AI Safety Report highlight research on dialect prejudice, including our work on covert racism in LLMs! www.nature.com/articles/s41... 161
Reposted by Valentin HofmannPaul Röttger @paul-rottger.bsky.social · 21/01/2025Today, we are releasing MSTS, a new Multimodal Safety Test Suite for vision-language models! MSTS is exciting because it tests for safety risks *created by multimodality*. Each prompt consists of a text + image that *only in combination* reveal their full unsafe meaning. 🧵 23011
Valentin Hofmann @valentinhofmann.bsky.social · 05/12/2024📢 New paper 📢 What generalization mechanisms shape the language skills of LLMs? Prior work has claimed that LLMs learn language via rules. We revisit the question and find that superficially rule-like behavior of LLMs can be traced to underlying analogical processes. 🧵 24510