Vilém Zouhar @zouhar.bsky.social · 02/10/2026arxiv has the opportunity to do the funniest thing ever 0181
Vilém Zouhar @zouhar.bsky.social · 29/09/2026I'm on the faculty job market for Fall 2027 assistant professorship. I work on the science of AI/NLP evaluation, which is currently undergoing a crisis. Get in touch! My research statement is public: vilda.net 1144
Vilém Zouhar @zouhar.bsky.social · 28/09/2026Feynman's 1974 "cargo cult science" is more relevant than ever in today's research landscape. We need to learn how to not fool ourselves and not just imitate science on the superficial level. calteches.library.caltech.edu/51/2/CargoCu...calteches.library.caltech.eduCargo Cult Science 090
Reposted by Vilém ZouharNiyati Bafna @niyatibafna.bsky.social · 23/09/2026Then we prove this theorem. 131
Vilém Zouhar @zouhar.bsky.social · 14/09/2026I spent good part of this year thinking about the methods of evaluation, specifically the way they're used at scale such as WMT or IWSLT. Happy to announce that they will be presented at different places at EMNLP: - Pearmut - Dynamic Annotation Allocation - cESA Annotation Protocol 191
Reposted by Vilém ZouharJannis Vamvas @vamvas.bsky.social · 07/09/2026Eine deutsche KI behauptet, Schweizerdeutsch zu können. Unser neuer Test sagt: «Chabis». vamvas.ch/benchmark-sw...vamvas.chEine deutsche KI behauptet, Schweizerdeutsch zu können. Unser neuer Test sagt: «Chabis» 152
Vilém Zouhar @zouhar.bsky.social · 04/09/2026Machine translation is not solved and it will take a while for it to be done arxiv.org/abs/2609.04173arxiv.orgLast Translation BenchmarkFor scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, standard benchmarks for ... 34612
Vilém Zouhar @zouhar.bsky.social · 22/08/2026Antropic blundered on the PR of destructivelly scanning books. Now everyone is imagining they're chopping the first prints of Mrs Dalloway but it's probably mostly "Step by Step Microsoft Access 2003". They could've gone with "We're saving books from the landfill" 0100
Vilém Zouhar @zouhar.bsky.social · 27/07/2026Last chance (3 days!) to participate in the refreshed WMT evaluation shared task. Pick any subtask and submit an automatic metric that's aligned with human annotations of translation quality across many languages. www2.statmt.org/wmt26/mteval... 121
Vilém Zouhar @zouhar.bsky.social · 21/07/2026definitely a good time to be shorting Jacobians this week 160
Vilém Zouhar @zouhar.bsky.social · 21/07/2026Several people on socials with no technical background told me they just retweet all papers stating they don't understand what the work is about. Asked about how they could trust the paper. They said when the paper is implemented at Google, the engineers will catch the errors. 082
Vilém Zouhar @zouhar.bsky.social · 20/07/2026Several reviewers with no technical background told me they reviewed for NeurIPS using agents, stating they didn't understand what the work is about. Asked how they could trust the review. They said when the paper hits the socials, the people will catch the errors. 1246
Vilém Zouhar @zouhar.bsky.social · 17/07/2026Not many people know this but this button was not invented to complain that reviewers didn't give you a higher score. 220
Vilém Zouhar @zouhar.bsky.social · 17/07/2026There are many things machine translation still can't do. Help us steer the next direction by contributing hard-to-translate inputs (and be on a cool paper). 2148
Vilém Zouhar @zouhar.bsky.social · 15/07/2026The best rebuttal I read is just 4 sentences in 3 bullet points clarifying my listed weaknesses. 1130
Vilém Zouhar @zouhar.bsky.social · 06/07/2026All my homies left for ACL/ICML and left me home managing a project *for which we're looking for coauthors* 🥸. Join us (also talk to @sethjsa.bsky.social @onadegibert.bsky.social @niyatibafna.bsky.social @patuchen.bsky.social @michellewastl.bsky.social @ayukh.bsky.social) 2104
Vilém Zouhar @zouhar.bsky.social · 02/07/2026The ironic twist is that ARR *did* save us by solving our obsession with publishing papers. I know at least one person who deferred from publishing her paper because of all the annoyance&hurdles in reviewing/ACing. 110
Vilém Zouhar @zouhar.bsky.social · 30/06/2026Be the reviewer you want (or your AC wants) to have. (The 1 goes to the AI written paper. No thank you for wasting 2 hours of my life. Pleasure reviewing the rest.) 030
Vilém Zouhar @zouhar.bsky.social · 30/06/2026me: spending 6 hours checking proofs in a paper I'm reviewing someone reviewing my paper: the paper does not evaluate whether its existence increases human evaluation adoption in practice, overall 2.5 170
Vilém Zouhar @zouhar.bsky.social · 28/06/2026The issue with Typst is that it's so much better than LaTeX, same as why Java was winning in corporate world. It's annoying to do things in LaTeX that go beyond the basics (text, basic styling, images, figures, tables). As a result, all papers and their tex code look the same. 171
Vilém Zouhar @zouhar.bsky.social · 28/06/2026Beginning to think that the reviewing disaster in comptuer science is caused by us not even liking to read papers. Some parents pay kids 1$ for each finished book. I propose we give each researcher +0.1 citations for successfully reading a paper. 180
Vilém Zouhar @zouhar.bsky.social · 16/06/2026We are at @eamt2026.bsky.social with a large-scale stealth project called "[redacted] Translation [redacted]". Looking for contributors across all languages to challenge MTs. Talk to @patuchen.bsky.social, @sethjsa.bsky.social and me to be a coauthor! 4123
Vilém Zouhar @zouhar.bsky.social · 08/06/2026We'll be hosting a tutorial at EAMT (already next week!), KONVENS and MT Marathon on human evaluation. Come learn with us! With support of @maikezufle.bsky.social and @patuchen.bsky.social 182
Vilém Zouhar @zouhar.bsky.social · 15/05/2026I reviewed for ICML and all I got was this lousy registration. 160
Vilém Zouhar @zouhar.bsky.social · 30/04/2026Deadline extension! - The task is simple: get audio + its translation and estimate how good it is. - Mark your name as the winner of the first Speech Translation Metrics Shared Task at IWSLT 2026 🏆 Predictions submission: May 7, 2026 Description paper: May 10, 2026 122
Vilém Zouhar @zouhar.bsky.social · 05/04/2026I love halucinated citations in papers. They serve as an obvious canary to AI-written papers. Without them, it takes a while to notice the discourse in writing doesn't make sense or that the science is shallow or unsound. 370
Vilém Zouhar @zouhar.bsky.social · 28/03/2026Come to La Palmaraie (EACL) for the first Multilingual Multicultural Evaluation workshop! 🧐 now. Organized by @pinzhen.bsky.social @hanxuhu.bsky.social @simi97k.bsky.social Wenhao Zhu @bazril.bsky.social Alexandra Birch @afaji.bsky.social Rico Sennrich @sarahooker.bsky.social 093
Vilém Zouhar @zouhar.bsky.social · 27/03/2026saddest conversation at eacl: - what do you work on? - mathematical modelling of evaluation - oh what kind of LLM is that? 080
Vilém Zouhar @zouhar.bsky.social · 26/03/2026all conference attendees under the age of 12 agree that playing subway surfer vastly improves the poster presentation experience 0100
Vilém Zouhar @zouhar.bsky.social · 12/03/2026Machine translation is tough to evaluate, partly because most of what you throw at is too easy. That doesn't at all mean that translation is solved; we're just not doing a good job finding interesting inputs. 1101
Vilém Zouhar @zouhar.bsky.social · 09/02/2026Quality estimation (automated metrics) are amazing. Truly. We would like to use them everywhere. That gets compute-expensive very quickly. We also don't know when they don't know. In "Early-Exit and Instant Confidence Translation Quality Estimation" (at EACL26) we fix that. 1153
Vilém Zouhar @zouhar.bsky.social · 28/01/2026How often is human evaluation skipped in papers/workflows just because it's too difficult to set up? Yet even small humeval can give so much more signal than automatic metrics. Introducing Pearmut, Human Evaluation of Translation Made Trivial🍐 arxiv.org/pdf/2601.02933 1180
Vilém Zouhar @zouhar.bsky.social · 14/01/2026Have you ever wondered how speech translation gets evaluated? Sadly, most speech evaluation downgrades to text-based metrics. Let’s do better! At IWSLT 2026, we’re launching the first-ever ✨Speech Translation Metrics Shared Task ✨! 181
Vilém Zouhar @zouhar.bsky.social · 03/01/2026Dissatisfied with EACL paper decisions? Fret not and submit your paper with ARR reviews to Multilingual Multicultural Evaluation workshop at EACL (both archival or nonarchival) until January 5th. 🔍🙂 multilingual-multicultural-evaluation.github.io 030
Reposted by Vilém ZouharGabriele Sarti @gsarti.com · 16/12/2025Now onwards to making language models transparent and trustworthy for everyone! 🚀 For those curious to know more about my thesis: - Web-optimized version: gsarti.com/phd-thesis/ - PDF: research.rug.nl/en/publicati... - Steal my Quarto template: github.com/gsarti/phd-t...gsarti.comFrom Insights to ImpactPh.D. Thesis, Center for Language and Cognition (CLCG), University of Groningen 0102
Vilém Zouhar @zouhar.bsky.social · 10/12/2025Do you have work on resources, metrics & methodologies for evaluating multilingual systems? Share it at the MME workshop 🕵️ co-located at EACL. Direct submission deadline in 10 days (December 19th)! multilingual-multicultural-evaluation.github.iomultilingual-multicultural-evaluation.github.ioMultilingual Multicultural Evaluation WorkshopLLMs in every language? Prove it. Showcase your work on rigorous, efficient, scalable, culture-aware multilingual benchmarking. 071
Vilém Zouhar @zouhar.bsky.social · 28/10/2025Let's talk about eval (automatic or human) and multilinguality at #EMNLP in Suzhou! 🇨🇳 - Efficient evaluation (Nov 5, 16:30, poster session 3) - MT difficulty (Nov 7, 12:30, findings 3) - COMET-poly (Nov 8, 11:00, WMT) (DM to meet 🌿 ) 4182
Vilém Zouhar @zouhar.bsky.social · 24/10/2025Grateful to receive the Google PhD Fellowship in NLP! 🙂 I am not secretive about having applied to 4 similar fellowships during my PhD before and not succeeding. Still, refining my research statement (part of the application) helped me tremendously in finding out the... inf.ethz.ch/news-and-eve...inf.ethz.chGoogle PhD Fellowships 2025Yutong Chen, Benedict Schlüter and Vilém Zouhar, all three of them doctoral students at the Department of Computer Science, have been awarded the Google PhD Fellowship. The programme was created to re... 1140
Vilém Zouhar @zouhar.bsky.social · 20/10/2025📢 Announcing the First Workshop on Multilingual and Multicultural Evaluation (MME) at #EACL2026 🇲🇦 MME focuses on resources, metrics & methodologies for evaluating multilingual systems! multilingual-multicultural-evaluation.github.io 📅 Workshop Mar 24–29, 2026 🗓️ Submit by Dec 19, 2025 13415
Vilém Zouhar @zouhar.bsky.social · 16/09/2025My two biggest take-aways are: - Standard testsets are too easy (Figure 1). - We can make testsets that are not easy (Figure 2). 😎 1173
Reposted by Vilém ZouharTom Kocmi @kocmitom.bsky.social · 23/08/2025We saw increased momentum in participation growth this year: 36 unique teams competing to improve the performance of MT. Furthermore, we added collected outputs of 24 popular LLMs and online systems. Reaching 50 evaluated systems in our annual benchmark. 131
Vilém Zouhar @zouhar.bsky.social · 25/07/2025The 2025 MT Evaluation shared task brings together the strengths of the previous Metrics and Quality Estimation tasks under a single, unified evaluation framework. The following tasks are now open for participants (deadline July 31st but participation has never been easier 🙂 ): 112
Vilém Zouhar @zouhar.bsky.social · 15/07/2025You have a budget to human-evaluate 100 inputs to your models, but your dataset is 10,000 inputs. Do not just pick 100 randomly!🙅 We can do better. "How to Select Datapoints for Efficient Human Evaluation of NLG Models?" shows how.🕵️ (random is still a devilishly good baseline) 2333
Vilém Zouhar @zouhar.bsky.social · 09/07/2025TIL that since python3.4 there's default `statistics` module with things like mean, mode, quantiles, variance, covariance, correlations, zscore, and more!. No more needless numpy imports! 250
Vilém Zouhar @zouhar.bsky.social · 07/07/2025Past iterations of the Terminology Shared Task don't come anywhere near the data quality and evaluation scrutiny of this one. In the era of LLM-as-MTs, participation has never been easier! 000
Vilém Zouhar @zouhar.bsky.social · 03/07/2025For the longest time I've been using Google Translate as a gateway to explain machine translation concepts to people as it's a tool that everyone knows. Now I get to contribute over the summer. 🌞 If you're near Mountain View, let's talk evaluation. 📏 0151
Vilém Zouhar @zouhar.bsky.social · 31/05/2025arxiv submission process got an update! (still requires a manual bbl) 4211