Sign in

Vilém Zouhar

@zouhar.bsky.social
3.7K followers 1.5K following 386 posts

PhD @ ETH Zürich | working on (multilingual) evaluation of NLP | on the academic job market | go #vegan | vilda.net

PostsRepliesMedia
Vilém Zouhar @zouhar.bsky.social · 02/10/2026
arxiv has the opportunity to do the funniest thing ever
0181
Vilém Zouhar @zouhar.bsky.social · 29/09/2026
I'm on the faculty job market for Fall 2027 assistant professorship. I work on the science of AI/NLP evaluation, which is currently undergoing a crisis. Get in touch! My research statement is public: vilda.net
1144
Vilém Zouhar @zouhar.bsky.social · 28/09/2026
Feynman's 1974 "cargo cult science" is more relevant than ever in today's research landscape. We need to learn how to not fool ourselves and not just imitate science on the superficial level. calteches.library.caltech.edu/51/2/CargoCu...
calteches.library.caltech.edu
Cargo Cult Science
090
Reposted by Vilém Zouhar
Niyati Bafna @niyatibafna.bsky.social · 23/09/2026
Then we prove this theorem.
131
Vilém Zouhar @zouhar.bsky.social · 14/09/2026
I spent good part of this year thinking about the methods of evaluation, specifically the way they're used at scale such as WMT or IWSLT. Happy to announce that they will be presented at different places at EMNLP: - Pearmut - Dynamic Annotation Allocation - cESA Annotation Protocol
191
Vilém Zouhar @zouhar.bsky.social · 08/09/2026
GPT-6, I thought you were AGI?
190
Reposted by Vilém Zouhar
Jannis Vamvas @vamvas.bsky.social · 07/09/2026
Eine deutsche KI behauptet, Schweizerdeutsch zu können. Unser neuer Test sagt: «Chabis». vamvas.ch/benchmark-sw...
vamvas.ch
Eine deutsche KI behauptet, Schweizerdeutsch zu können. Unser neuer Test sagt: «Chabis»
152
Vilém Zouhar @zouhar.bsky.social · 04/09/2026
Machine translation is not solved and it will take a while for it to be done arxiv.org/abs/2609.04173
arxiv.org
Last Translation Benchmark
For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, standard benchmarks for ...
34612
Vilém Zouhar @zouhar.bsky.social · 22/08/2026
Antropic blundered on the PR of destructivelly scanning books. Now everyone is imagining they're chopping the first prints of Mrs Dalloway but it's probably mostly "Step by Step Microsoft Access 2003". They could've gone with "We're saving books from the landfill"
0100
Vilém Zouhar @zouhar.bsky.social · 27/07/2026
Last chance (3 days!) to participate in the refreshed WMT evaluation shared task. Pick any subtask and submit an automatic metric that's aligned with human annotations of translation quality across many languages. www2.statmt.org/wmt26/mteval...
121
Vilém Zouhar @zouhar.bsky.social · 21/07/2026
definitely a good time to be shorting Jacobians this week
160
Vilém Zouhar @zouhar.bsky.social · 21/07/2026
Several people on socials with no technical background told me they just retweet all papers stating they don't understand what the work is about. Asked about how they could trust the paper. They said when the paper is implemented at Google, the engineers will catch the errors.
082
Vilém Zouhar @zouhar.bsky.social · 20/07/2026
Several reviewers with no technical background told me they reviewed for NeurIPS using agents, stating they didn't understand what the work is about. Asked how they could trust the review. They said when the paper hits the socials, the people will catch the errors.
1246
Vilém Zouhar @zouhar.bsky.social · 17/07/2026
Not many people know this but this button was not invented to complain that reviewers didn't give you a higher score.
220
Vilém Zouhar @zouhar.bsky.social · 17/07/2026
There are many things machine translation still can't do. Help us steer the next direction by contributing hard-to-translate inputs (and be on a cool paper).
2148
Vilém Zouhar @zouhar.bsky.social · 15/07/2026
The best rebuttal I read is just 4 sentences in 3 bullet points clarifying my listed weaknesses.
1130
Vilém Zouhar @zouhar.bsky.social · 06/07/2026
All my homies left for ACL/ICML and left me home managing a project *for which we're looking for coauthors* 🥸. Join us (also talk to @sethjsa.bsky.social @onadegibert.bsky.social @niyatibafna.bsky.social @patuchen.bsky.social @michellewastl.bsky.social @ayukh.bsky.social)
2104
Vilém Zouhar @zouhar.bsky.social · 02/07/2026
The ironic twist is that ARR *did* save us by solving our obsession with publishing papers. I know at least one person who deferred from publishing her paper because of all the annoyance&hurdles in reviewing/ACing.
110
Vilém Zouhar @zouhar.bsky.social · 30/06/2026
Be the reviewer you want (or your AC wants) to have. (The 1 goes to the AI written paper. No thank you for wasting 2 hours of my life. Pleasure reviewing the rest.)
030
Vilém Zouhar @zouhar.bsky.social · 30/06/2026
me: spending 6 hours checking proofs in a paper I'm reviewing someone reviewing my paper: the paper does not evaluate whether its existence increases human evaluation adoption in practice, overall 2.5
170
Vilém Zouhar @zouhar.bsky.social · 28/06/2026
The issue with Typst is that it's so much better than LaTeX, same as why Java was winning in corporate world. It's annoying to do things in LaTeX that go beyond the basics (text, basic styling, images, figures, tables). As a result, all papers and their tex code look the same.
171
Vilém Zouhar @zouhar.bsky.social · 28/06/2026
Beginning to think that the reviewing disaster in comptuer science is caused by us not even liking to read papers. Some parents pay kids 1$ for each finished book. I propose we give each researcher +0.1 citations for successfully reading a paper.
180
Vilém Zouhar @zouhar.bsky.social · 16/06/2026
We are at @eamt2026.bsky.social with a large-scale stealth project called "[redacted] Translation [redacted]". Looking for contributors across all languages to challenge MTs. Talk to @patuchen.bsky.social, @sethjsa.bsky.social and me to be a coauthor!
4123
Vilém Zouhar @zouhar.bsky.social · 08/06/2026
We'll be hosting a tutorial at EAMT (already next week!), KONVENS and MT Marathon on human evaluation. Come learn with us! With support of @maikezufle.bsky.social and @patuchen.bsky.social
182
Vilém Zouhar @zouhar.bsky.social · 15/05/2026
I reviewed for ICML and all I got was this lousy registration.
160
Vilém Zouhar @zouhar.bsky.social · 14/05/2026
environmental storytelling
0130
Vilém Zouhar @zouhar.bsky.social · 30/04/2026
Deadline extension! - The task is simple: get audio + its translation and estimate how good it is. - Mark your name as the winner of the first Speech Translation Metrics Shared Task at IWSLT 2026 🏆 Predictions submission: May 7, 2026 Description paper: May 10, 2026
122
Vilém Zouhar @zouhar.bsky.social · 05/04/2026
I love halucinated citations in papers. They serve as an obvious canary to AI-written papers. Without them, it takes a while to notice the discourse in writing doesn't make sense or that the science is shallow or unsound.
370
Vilém Zouhar @zouhar.bsky.social · 28/03/2026
Come to La Palmaraie (EACL) for the first Multilingual Multicultural Evaluation workshop! 🧐 now. Organized by @pinzhen.bsky.social @hanxuhu.bsky.social @simi97k.bsky.social Wenhao Zhu @bazril.bsky.social Alexandra Birch @afaji.bsky.social Rico Sennrich @sarahooker.bsky.social
093
Vilém Zouhar @zouhar.bsky.social · 27/03/2026
saddest conversation at eacl: - what do you work on? - mathematical modelling of evaluation - oh what kind of LLM is that?
080
Vilém Zouhar @zouhar.bsky.social · 26/03/2026
all conference attendees under the age of 12 agree that playing subway surfer vastly improves the poster presentation experience
0100
Vilém Zouhar @zouhar.bsky.social · 12/03/2026
Machine translation is tough to evaluate, partly because most of what you throw at is too easy. That doesn't at all mean that translation is solved; we're just not doing a good job finding interesting inputs.
1101
Vilém Zouhar @zouhar.bsky.social · 09/02/2026
Quality estimation (automated metrics) are amazing. Truly. We would like to use them everywhere. That gets compute-expensive very quickly. We also don't know when they don't know. In "Early-Exit and Instant Confidence Translation Quality Estimation" (at EACL26) we fix that.
1153
Vilém Zouhar @zouhar.bsky.social · 28/01/2026
How often is human evaluation skipped in papers/workflows just because it's too difficult to set up? Yet even small humeval can give so much more signal than automatic metrics. Introducing Pearmut, Human Evaluation of Translation Made Trivial🍐 arxiv.org/pdf/2601.02933
1180
Vilém Zouhar @zouhar.bsky.social · 14/01/2026
Have you ever wondered how speech translation gets evaluated? Sadly, most speech evaluation downgrades to text-based metrics. Let’s do better! At IWSLT 2026, we’re launching the first-ever ✨Speech Translation Metrics Shared Task ✨!
181
Vilém Zouhar @zouhar.bsky.social · 03/01/2026
Dissatisfied with EACL paper decisions? Fret not and submit your paper with ARR reviews to Multilingual Multicultural Evaluation workshop at EACL (both archival or nonarchival) until January 5th. 🔍🙂 multilingual-multicultural-evaluation.github.io
030
Reposted by Vilém Zouhar
Gabriele Sarti @gsarti.com · 16/12/2025
Now onwards to making language models transparent and trustworthy for everyone! 🚀 For those curious to know more about my thesis: - Web-optimized version: gsarti.com/phd-thesis/ - PDF: research.rug.nl/en/publicati... - Steal my Quarto template: github.com/gsarti/phd-t...
gsarti.com
From Insights to Impact
Ph.D. Thesis, Center for Language and Cognition (CLCG), University of Groningen
0102
Vilém Zouhar @zouhar.bsky.social · 10/12/2025
Do you have work on resources, metrics & methodologies for evaluating multilingual systems? Share it at the MME workshop 🕵️ co-located at EACL. Direct submission deadline in 10 days (December 19th)! multilingual-multicultural-evaluation.github.io
multilingual-multicultural-evaluation.github.io
Multilingual Multicultural Evaluation Workshop
LLMs in every language? Prove it. Showcase your work on rigorous, efficient, scalable, culture-aware multilingual benchmarking.
071
Vilém Zouhar @zouhar.bsky.social · 28/10/2025
Let's talk about eval (automatic or human) and multilinguality at #EMNLP in Suzhou! 🇨🇳 - Efficient evaluation (Nov 5, 16:30, poster session 3) - MT difficulty (Nov 7, 12:30, findings 3) - COMET-poly (Nov 8, 11:00, WMT) (DM to meet 🌿 )
4182
Vilém Zouhar @zouhar.bsky.social · 24/10/2025
Grateful to receive the Google PhD Fellowship in NLP! 🙂 I am not secretive about having applied to 4 similar fellowships during my PhD before and not succeeding. Still, refining my research statement (part of the application) helped me tremendously in finding out the... inf.ethz.ch/news-and-eve...
inf.ethz.ch
Google PhD Fellowships 2025
Yutong Chen, Benedict Schlüter and Vilém Zouhar, all three of them doctoral students at the Department of Computer Science, have been awarded the Google PhD Fellowship. The programme was created to re...
1140
Vilém Zouhar @zouhar.bsky.social · 20/10/2025
📢 Announcing the First Workshop on Multilingual and Multicultural Evaluation (MME) at #EACL2026 🇲🇦 MME focuses on resources, metrics & methodologies for evaluating multilingual systems! multilingual-multicultural-evaluation.github.io 📅 Workshop Mar 24–29, 2026 🗓️ Submit by Dec 19, 2025
13415
Vilém Zouhar @zouhar.bsky.social · 16/09/2025
My two biggest take-aways are: - Standard testsets are too easy (Figure 1). - We can make testsets that are not easy (Figure 2). 😎
1173
Reposted by Vilém Zouhar
Tom Kocmi @kocmitom.bsky.social · 23/08/2025
We saw increased momentum in participation growth this year: 36 unique teams competing to improve the performance of MT. Furthermore, we added collected outputs of 24 popular LLMs and online systems. Reaching 50 evaluated systems in our annual benchmark.
131
Vilém Zouhar @zouhar.bsky.social · 25/07/2025
The 2025 MT Evaluation shared task brings together the strengths of the previous Metrics and Quality Estimation tasks under a single, unified evaluation framework. The following tasks are now open for participants (deadline July 31st but participation has never been easier 🙂 ):
112
Vilém Zouhar @zouhar.bsky.social · 15/07/2025
You have a budget to human-evaluate 100 inputs to your models, but your dataset is 10,000 inputs. Do not just pick 100 randomly!🙅 We can do better. "How to Select Datapoints for Efficient Human Evaluation of NLG Models?" shows how.🕵️ (random is still a devilishly good baseline)
2333
Vilém Zouhar @zouhar.bsky.social · 09/07/2025
TIL that since python3.4 there's default `statistics` module with things like mean, mode, quantiles, variance, covariance, correlations, zscore, and more!. No more needless numpy imports!
250
Vilém Zouhar @zouhar.bsky.social · 07/07/2025
Past iterations of the Terminology Shared Task don't come anywhere near the data quality and evaluation scrutiny of this one. In the era of LLM-as-MTs, participation has never been easier!
000
Vilém Zouhar @zouhar.bsky.social · 03/07/2025
Thank you for your response. I will keep my score.
2190
Vilém Zouhar @zouhar.bsky.social · 03/07/2025
For the longest time I've been using Google Translate as a gateway to explain machine translation concepts to people as it's a tool that everyone knows. Now I get to contribute over the summer. 🌞 If you're near Mountain View, let's talk evaluation. 📏
0151
Vilém Zouhar @zouhar.bsky.social · 31/05/2025
arxiv submission process got an update! (still requires a manual bbl)
4211