Sign in

Jindřich Libovický

@jlibovicky.bsky.social
575 followers 260 following 62 posts

Researcher at Charles University | multilingual natural language processing, machine translation

PostsRepliesMedia
Jindřich Libovický @jlibovicky.bsky.social · 01/10/2026
TokShop is just one week away! 🎉 I'm proud of the lineup of invited speakers we've put together. Meet them here: tokenization-workshop.github.io/speakers
tokenization-workshop.github.io
Second Tokenization Workshop @ COLM 2026 | Speakers
010
Jindřich Libovický @jlibovicky.bsky.social · 01/10/2026
Everything you always wanted to know about tokenization but were afraid to ask! 🧩 @mcognetta.bsky.social l and @uvp.bsky.social nagged 30+ researchers (myself included) until we ended up with the most comprehensive survey on tokenization in NLP: www.alphaxiv.org/abs/2609.tok...
alphaxiv.org
Tokenization: A Survey for Modern NLP
Tokenization is presented as a core language-model design choice that shapes sequence length, computational cost, multilingual equity, evaluation, and security—not merely as preprocessing. The...
191
Jindřich Libovický @jlibovicky.bsky.social · 30/09/2026
ACL announcing the deadline for ACL 2027...
031
Jindřich Libovický @jlibovicky.bsky.social · 21/09/2026
I really enjoyed giving a talk at the Machine Translation Marathon 2026 about our research on tokenization evaluation. Thank you for having me! For those interested, the slides from my presentation are available here: docs.google.com/presentation...
072
Jindřich Libovický @jlibovicky.bsky.social · 03/09/2026
Yes. It's definitely because of the training, plus perhaps an unknown inductive bias in the architecture. We sort of expected that, for Qwen, Chinese might serve as the representation pivot, but that was not the case.
110
Jindřich Libovický @jlibovicky.bsky.social · 02/09/2026
Answer: similarity of that language's representation to English. Unless two languages are very close (e.g., Czech/Slovak), alignment with English matters much more than similarity to the source language. Want to know more? Catch Adnan at #EMNLP2026 in Budapest!
120
Jindřich Libovický @jlibovicky.bsky.social · 02/09/2026
Very happy to share that Adnan Al Ali's first big PhD project has been accepted to #EMNLP2026! 🎉 Joint work with @kathaem.bsky.social, me & Alex Fraser. Paper: arxiv.org/abs/2608.03446 What predicts multilingual LLM performance in a language? 🧵
arxiv.org
Predicting Multilingual Classification and Translation Performance of LLMs with Cross-Lingual Alignment -- Is English Enough?
Multilingual large language models (LLMs) have been shown to perform better on non-English classification tasks when the representations of the given language are more aligned to English within the mo...
161
Jindřich Libovický @jlibovicky.bsky.social · 07/07/2026
It probably is. But human annotation is expensive, and some guidance (even based on LLMs) on what metric to use is better than no guidance and doing LLM as a judge with the first prompt that comes to your mind.
010
Jindřich Libovický @jlibovicky.bsky.social · 07/07/2026
Very proud of my student Lukáš Eigler presenting this at the ACL Student Research Workshop. TL;DR: you can validate NLP evaluation metrics with synthetic LLM judgments — the rankings track human ones almost perfectly. 📄 aclanthology.org/2026.acl-srw.125
aclanthology.org
LLM as a Meta-Judge: Synthetic Data for NLP Evaluation Metric Validation
Lukáš Eigler, Jindřich Libovický, David Hurych. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop). 2026.
110
Jindřich Libovický @jlibovicky.bsky.social · 06/07/2026
Congrats Dušan and @abyste.bsky.social ! 🎉 Btw. the demo ships pre-loaded with FLORES-200, so you can compare 21 tokenizers across 222 languages right out of the box. 📄 aclanthology.org/2026.acl-demo.41
aclanthology.org
TokCollate: A Comprehensive Tool for Tokenizer Evaluation and Visualization across Languages
Dušan Variš, Abishek Stephen, Jindřich Libovický. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations). 2026.
000
Jindřich Libovický @jlibovicky.bsky.social · 06/07/2026
Dušan made something like TensorBoard for tokenization 📊 Token length, entropy, vocab overlap, JS divergence, alignment scores: sliced by language family, script, region, speakers, data availability. 🔗 Code: github.com/ufal/TokCollate 🔗 Demo: quest.ms.mff.cuni.cz/tokcollate
140
Jindřich Libovický @jlibovicky.bsky.social · 02/07/2026
AC observation this ARR cycle: I'm getting way more reviews and way earlier, than usual. I suspect it's the draconian desk-reject threats working, and not a sign of better reviewing culture. I worry that this gets read as 'threats work, do more of them, discipline the lazy researchers'.
020
Reposted by Jindřich Libovický
Tokenization Workshop (TokShop) @COLM2026 @tokshop.bsky.social · 23/06/2026
📣 TokShop 2026 deadline extended! 🗓️ New submission deadline: Friday, June 26, 2026 (AoE) Research papers (up to 9 pages) and extended abstracts (up to 2 pages) are welcome. Submit: openreview.net/group?id=col... More info: tokenization-workshop.github.io
034
Jindřich Libovický @jlibovicky.bsky.social · 15/06/2026
This dataset is a part of master thesis of my student Michal Tichý, who just defended his thesis! 🎉 Congrats!
000
Jindřich Libovický @jlibovicky.bsky.social · 15/06/2026
Is "Je to obrovský problém, tvrdí české ministerstvo obrany." Czech or Slovak? Both. Language ID systems must pick one: and that's just one way they break. Our new benchmark CHALIS exposes how fragile they really are. 📄 arxiv.org/abs/2606.06088 🤗 huggingface.co/datasets/michal-tichy/CHALIS
arxiv.org
CHALIS: A Challenge Dataset for Language Identification in Difficult Scenarios
We present CHALIS (Challenging Language Identification Samples), a new benchmark dataset explicitly designed to address difficult cases in language identification: cousin languages and orthographic no...
150
Jindřich Libovický @jlibovicky.bsky.social · 09/06/2026
Lukáš Eigler defended his thesis (co-supervised with David Hurych, @valeoai.bsky.social) 🎉 Congrats! #NLP metric validation needs 🐌💰 human judgment data. Our fix: generate synthetic data for metric validation. ✅ Tested on MT, QA, summarization. To appear #ACL2026 SRW: arxiv.org/abs/2603.09403
arxiv.org
LLM as a Meta-Judge: Synthetic Data for NLP Evaluation Metric Validation
Validating evaluation metrics for NLG typically relies on expensive and time-consuming human annotations, which predominantly exist only for English datasets. We propose LLM as a Meta-Judge, a scalabl...
091
Jindřich Libovický @jlibovicky.bsky.social · 05/06/2026
I have to grade 200+ ML exams/year. To survive, I coded a bunch of tools: unique exam generation, handwriting recognition, and per-question stats. Check out the details on my blog. 📊 jlibovicky.github.io/2026/06/04/G...
jlibovicky.github.io
How I grade 200 exams every year
Three years ago, I took the introductory machine learning course over from Milan Straka, and one of the problems I had to deal with was: how do I grade 250 written exams without it consuming my entire...
160
Jindřich Libovický @jlibovicky.bsky.social · 22/05/2026
Send your work to TokShop! tokenization-workshop.github.io
010
Jindřich Libovický @jlibovicky.bsky.social · 14/05/2026
TokShop is happening again!
040
Jindřich Libovický @jlibovicky.bsky.social · 30/03/2026
Nice work by my student @gianlucavico.bsky.social on a topic close to home: crowdsourcing Piedmontese to test LLMs on non-standard orthography. New dataset covering tokenization, classification & translation.
041
Reposted by Jindřich Libovický
Institute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 28/03/2026
On the Credibility of Evaluating LLMs using Survey Questions by @jlibovicky.bsky.social aclanthology.org/2026.mme-mai... Survey-based LLM value evals are unreliable — prompting & decoding choices skew results drastically.
172
Jindřich Libovický @jlibovicky.bsky.social · 26/03/2026
So proud of my new PhD student Adnan Al Ali for presenting their master's thesis work at EACL! 🎓 A great contribution to understanding bias in AI text detectors across languages.
020
Jindřich Libovický @jlibovicky.bsky.social · 26/03/2026
I'm at #EACL2026 in Rabat 🇲🇦. Find me and talk to me about tokenization or multilingual model eval. Also, check out our work on how to eval morphological plausibility of your tokenizer if you don't have gold segmentation data, but you happen to have morphosyntactic features 👇
140
Jindřich Libovický @jlibovicky.bsky.social · 20/03/2026
Thanks for organizing this! I am definitely coming, and everyone interested in tokenization at #EACL2025 should too. 🫵
010
Jindřich Libovický @jlibovicky.bsky.social · 09/03/2026
Intro to NLP for bachelor students ufal.mff.cuni.cz/courses/npfl... Most of it would be History of NLP in the Stanford slides 😀
ufal.mff.cuni.cz
Natural Language Processing | ÚFAL
001
Jindřich Libovický @jlibovicky.bsky.social · 09/03/2026
None of this would work without my TAs: Dušan Variš, Tomáš Musil, Jan Bronec, @gianlucavico.bsky.social , Adnan Al Ali, Kristýna Onderková, and @straka-milan.bsky.social taking care of ReCodEx: recodex.mff.cuni.cz. Thank you 🙏
recodex.mff.cuni.cz
030
Jindřich Libovický @jlibovicky.bsky.social · 09/03/2026
Some students find the assignments too time-consuming. Fair. But here's what the data shows over 3 years: 📉 Forum questions dropped ~4× 📈 Full bonus points: 20% → 27% → 52% 📉 Avg. test attempts: 2.7 → 2.4 → 1.9 Asking less, achieving more, iterating less. 🤔
140
Jindřich Libovický @jlibovicky.bsky.social · 09/03/2026
3rd run teaching ML to 250+ bachelor students (with great materials originaly by @straka-milan.bsky.social). Core philosophy: explain the math, implement algorithms from scratch, Kaggle-style competitions, all auto-graded. ufal.mff.cuni.cz/courses/npfl... But look what LLMs did to the course 👇
ufal.mff.cuni.cz
Introduction to Machine Learning with Python | ÚFAL
130
Jindřich Libovický @jlibovicky.bsky.social · 04/03/2026
Spent time making AI-generated images of Bayes' Rule, Laplace Smoothing, Markov Chains & Shannon Entropy for class today 🎨🤖 Even though the images are objectively hilarious, none of the 50 students in the room laughed. Or even smiled. 💀
151
Jindřich Libovický @jlibovicky.bsky.social · 16/02/2026
More interesting: humans are predictably inconsistent in their values. LLMs capture this but overgeneralize: they become more stereotypically consistent than actual humans. After several rejections, finally publishable. To appear at the Multilingual Multicultural Evaluation workshop at EACL 2026.
022
Jindřich Libovický @jlibovicky.bsky.social · 16/02/2026
I reviewed papers evaluating LLM values using sociology questionnaires. Different methods, different results. Didn't trust them, so I tested it myself. Methodology matters. Short answers vs CoT, squared err vs KL div.: each changes which populations an LLM "aligns" with. www.arxiv.org/pdf/2602.04033
arxiv.org
151
Jindřich Libovický @jlibovicky.bsky.social · 03/02/2026
We have updated the pre-print on CUS-QA, benchmark for regional knowledge about Czechia, Slovakia and Ukraine arxiv.org/abs/2507.22752 Now, there are results of retrieval-augmented generation and more detailed analysis of model performance depending on the topic of the question or visual context.
arxiv.org
CUS-QA: Local-Knowledge-Oriented Open-Ended Question Answering Dataset
We introduce CUS-QA, a benchmark for evaluation of open-ended regional question answering that encompasses both textual and visual modalities. We also provide strong baselines using state-of-the-art l...
071
Jindřich Libovický @jlibovicky.bsky.social · 02/02/2026
👉 What do we do? We use the good old IBM1 model to align subwords with morphological features from Unimorph and we show it captures the same thing as morpheme boundary recall. 👉 Why it matters? For many languages good segmentation data is missing. Morphological features are more widely available.
041
Jindřich Libovický @jlibovicky.bsky.social · 02/02/2026
We (= mostly @abyste.bsky.social) developed a way to evaluate how morphological a #tokenization is w/o gold segmentation labels. arxiv.org/abs/2601.18536 The key: align subword tokens with morphological features from UniMorph using IBM Model 1. To appear in EACL 2026 Findings.
arxiv.org
Evaluating Morphological Plausibility of Subword Tokenization via Statistical Alignment with Morpho-Syntactic Features
We present a novel metric for the evaluation of the morphological plausibility of subword segmentation. Unlike the typically used morpheme boundary or retrieval F-score, which requires gold segmentati...
191
Jindřich Libovický @jlibovicky.bsky.social · 23/12/2025
Happy holidays! 🎄🎅🤩🎁
040
Jindřich Libovický @jlibovicky.bsky.social · 10/11/2025
Attenzione! 🇮🇹 Know Piedmontese or Neapolitan speakers? @gianlucavico.bsky.social is collecting crowd-sourced translations to evaluate LLM performance on these regional languages. Partecipate!
021
Jindřich Libovický @jlibovicky.bsky.social · 21/10/2025
Cultural awareness is trickier. Different data for different cultures means we can't really compare performance across cultures in a straightforward way. And there's no clear optimization target for cultural awareness beyond curating diverse training data.
010
Jindřich Libovický @jlibovicky.bsky.social · 21/10/2025
☝️🧵 Most current approaches emphasize langauge neutrality: about two-thirds of VL benchmarks use translation-based evaluation. This makes sense because we can explicitly train for language neutrality when we have parallel data. But... 🧵👇
100
Jindřich Libovický @jlibovicky.bsky.social · 21/10/2025
With @andrei-a-manea.bsky.social, we posted a survey on multilingual vision-language models 👉 arxiv.org/pdf/2509.22123 We reviewed 31 models+21 benchmarks. There's a tension between language neutrality (same results across languages) & cultural awareness (context matters differently across cultures)
arxiv.org
132
Jindřich Libovický @jlibovicky.bsky.social · 01/09/2025
Most vision-language models only work in English. We explore how different parallel data types (machine-translated vs authentic captions) affect cross-lingual transfer. Key finding: authentic data can outperform machine translation, and multilingual training beats bilingual approaches. #NLP
020
Jindřich Libovický @jlibovicky.bsky.social · 01/09/2025
So proud of my PhD student @andrei-a-manea.bsky.social for his first first-author publication! 🎉 He presented this work last week at TSD. Investigating the Effect of Parallel Data in the Cross-Lingual Transfer for Vision-Language Encoders arxiv.org/pdf/2504.21681
arxiv.org
160
Jindřich Libovický @jlibovicky.bsky.social · 25/08/2025
For evaluation researchers: Simple string-overlap metrics (BLEU, chrF) work surprisingly well for factual QA. 🤔 When answers are mostly named entities, exact matches matter more than we thought. LLM-as-judge 🦙🧑‍⚖️ correlates best with human judgment, though.
110
Jindřich Libovický @jlibovicky.bsky.social · 25/08/2025
The results are... humbling 😅 Even the best models: >40% accuracy on textual questions <30% on visual questions Often perform better in English than the local language (!!) Visual QA with regional images is especially challenging.
100
Jindřich Libovický @jlibovicky.bsky.social · 25/08/2025
The problem: Most QA benchmarks focus on globally known facts. But real users ask about local geography, culture, and history. We collected questions from native speakers in Czechia 🇨🇿, Slovakia 🇸🇰, and Ukraine 🇺🇦 about facts locals know but outsiders don't.
100
Jindřich Libovický @jlibovicky.bsky.social · 25/08/2025
🧵 We're releasing CUS-QA - a new benchmark for testing LLMs on regional knowledge! Find out what your model knows about Czechia 🇨🇿, Slovakia 🇸🇰, and Ukraine 🇺🇦! 👉 Textual and visual questions, answers, and human judgment on model outputs! huggingface.co/datasets/ufa... www.arxiv.org/abs/2507.22752
huggingface.co
ufal/cus-qa · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
1163
Jindřich Libovický @jlibovicky.bsky.social · 01/08/2025
Stay tuned, we will release the dataset soon...
020
Reposted by Jindřich Libovický
Jindra Helcl @jindrahelcl.bsky.social · 29/07/2025
We need to have poster fights at the end of every conference.
031
Jindřich Libovický @jlibovicky.bsky.social · 28/07/2025
Just presented MAGBIG, a new dataset and evaluation methodology for gender bias in multilingual text-to-image generation. Grammatical gender matters when studying these biases across languages! Thanks to Felix Friedrich, @kathaem.bsky.social and all co-authors - it was fun to work on this together!
020
Jindřich Libovický @jlibovicky.bsky.social · 27/07/2025
This week I am at #ACL2025NLP in Vienna 🎡🇦🇹. Find me 🕵️ or message 💌 me if you want to chat about multilinguality or tokenization. Stop 🛑 by our poster on gender bias in text-to-image generation on Monday aclanthology.org/2025.acl-lon...
070
Reposted by Jindřich Libovický
Tokenization Workshop (TokShop) @COLM2026 @tokshop.bsky.social · 02/06/2025
TokShop @ #ICML2025 got way more submissions than expected! 📈 We could really use a few more reviewers to help out. If you have the capacity to review a #tokenization paper by Saturday, please fill out this form: forms.gle/32A6sQHQrMSb... 🙏
forms.gle
TokShop 2025
Registering interest in all things tokenization at TokShop @ ICML 2025 (July 18) Consider joining the Google group for future updates! https://groups.google.com/g/tokshop
004