Sign in

Jindřich Libovický

@jlibovicky.bsky.social
576 followers 259 following 62 posts

Researcher at Charles University | multilingual natural language processing, machine translation

PostsRepliesMedia
Jindřich Libovický @jlibovicky.bsky.social · 30/09/2026
ACL announcing the deadline for ACL 2027...
031
Jindřich Libovický @jlibovicky.bsky.social · 21/09/2026
I really enjoyed giving a talk at the Machine Translation Marathon 2026 about our research on tokenization evaluation. Thank you for having me! For those interested, the slides from my presentation are available here: docs.google.com/presentation...
072
Jindřich Libovický @jlibovicky.bsky.social · 06/07/2026
Dušan made something like TensorBoard for tokenization 📊 Token length, entropy, vocab overlap, JS divergence, alignment scores: sliced by language family, script, region, speakers, data availability. 🔗 Code: github.com/ufal/TokCollate 🔗 Demo: quest.ms.mff.cuni.cz/tokcollate
140
Jindřich Libovický @jlibovicky.bsky.social · 09/03/2026
Some students find the assignments too time-consuming. Fair. But here's what the data shows over 3 years: 📉 Forum questions dropped ~4× 📈 Full bonus points: 20% → 27% → 52% 📉 Avg. test attempts: 2.7 → 2.4 → 1.9 Asking less, achieving more, iterating less. 🤔
140
Jindřich Libovický @jlibovicky.bsky.social · 04/03/2026
Spent time making AI-generated images of Bayes' Rule, Laplace Smoothing, Markov Chains & Shannon Entropy for class today 🎨🤖 Even though the images are objectively hilarious, none of the 50 students in the room laughed. Or even smiled. 💀
151
Jindřich Libovický @jlibovicky.bsky.social · 16/02/2026
More interesting: humans are predictably inconsistent in their values. LLMs capture this but overgeneralize: they become more stereotypically consistent than actual humans. After several rejections, finally publishable. To appear at the Multilingual Multicultural Evaluation workshop at EACL 2026.
022
Jindřich Libovický @jlibovicky.bsky.social · 02/02/2026
👉 What do we do? We use the good old IBM1 model to align subwords with morphological features from Unimorph and we show it captures the same thing as morpheme boundary recall. 👉 Why it matters? For many languages good segmentation data is missing. Morphological features are more widely available.
041
Jindřich Libovický @jlibovicky.bsky.social · 23/12/2025
Happy holidays! 🎄🎅🤩🎁
040
Jindřich Libovický @jlibovicky.bsky.social · 25/08/2025
The problem: Most QA benchmarks focus on globally known facts. But real users ask about local geography, culture, and history. We collected questions from native speakers in Czechia 🇨🇿, Slovakia 🇸🇰, and Ukraine 🇺🇦 about facts locals know but outsiders don't.
100
Jindřich Libovický @jlibovicky.bsky.social · 27/07/2025
This week I am at #ACL2025NLP in Vienna 🎡🇦🇹. Find me 🕵️ or message 💌 me if you want to chat about multilinguality or tokenization. Stop 🛑 by our poster on gender bias in text-to-image generation on Monday aclanthology.org/2025.acl-lon...
070
Jindřich Libovický @jlibovicky.bsky.social · 30/04/2025
Attending #NAACL2025 virtually. Since 2022, I've been training a classifier on papers I read to tackle the arXiv madness. Ran it on the NAACL proceedings for my personalized watch list. 🤓📺 However, it's far from perfect: Multilingual cultural awareness is great, but where is tokenization? 🤷
220
Jindřich Libovický @jlibovicky.bsky.social · 24/12/2024
Happy holidays! 🎄🎅🤩🎁
080