Reposted by Catherine ArnettMarco @mcognetta.bsky.social · 30/09/2026🚨 [Token][ization] Paper Alert 🚨 Tokenization is a wildly understudied area of language modeling despite it having effects across all of NLP. Over the past ~8 months, 32 (!) tokenizer researchers put together the most comprehensive survey of the field. Check it out! 112935
Reposted by Catherine ArnettLeshem (Legend) Choshen @EMNLP @lchoshen.bsky.social · 28/08/2026Will the same model become a different player depending on the interface language? We report our findings in our paper "Skill Issue: Are Skills Language-Invariant in LLMs?". And with it, TextArena is now multilingual 🌍 with 65 games in 193 languages. 🧵 1172
Catherine Arnett @catherinearnett.bsky.social · 28/08/2026Xiulin has been doing some extremely interesting and important work about how to go about actually comparable language model evaluation across models and across languages. We have some recommendations - check them out in our new preprint and catch the paper at #EMNLP2026! 1110
Catherine Arnett @catherinearnett.bsky.social · 20/08/2026As a personal update, I have moved to the city of Oxford where I’ll continue working from EleutherAI! Looking forward to meeting people in the area! 1150
Reposted by Catherine ArnettMultilingual Representation Workshop @ EMNLP 2026 @mrl-workshop.bsky.social · 28/07/2026We are super impressed by the submissions for Phase 1 of the MRL Shared Task! :) In order for us to thoroughly review the datasets, we will extend the deadline to August 14! The audio collection phase will now happen in September. Please stay tuned here (or on Discord) for more updates! 011
Reposted by Catherine ArnettJames Michaelov @jamichaelov.bsky.social · 11/06/2026Seems like a good time to share our new preprint about model openness! (with @catherinearnett.bsky.social @tylerachang.bsky.social Pamela D. Rivière, Samuel M. Taylor @camrobjones.bsky.social @seantrott.bsky.social @rplevy.bsky.social Ben Bergen, and Micah Altman): arxiv.org/abs/2603.26539 1193
Reposted by Catherine ArnettStella Biderman @stellaathena.bsky.social · 10/06/2026In film, "we'll fix it in post" is what you say when something went wrong on set and you don't want to redo it. AI research has made it our entire methodology: train the model, then patch whatever comes out. Our new ICML oral argues this can't be the basis of a science of AI. 🧵 310823
Catherine Arnett @catherinearnett.bsky.social · 08/06/2026This is a great opportunity for students and early-stage researchers. Please contribute if you can! 020
Catherine Arnett @catherinearnett.bsky.social · 02/06/2026The new and expanded version of Global PIQA is out with over twice as many items. Well done to all the contributors! 1113
Reposted by Catherine ArnettMultilingual Representation Workshop @ EMNLP 2026 @mrl-workshop.bsky.social · 27/05/2026📢 Call for Papers: 6th Multilingual Representation Learning Workshop at EMNLP in Budapest, Hungary! Join us and submit your works relating to multilingual NLP Speakers to be announced, so stay tuned! 👀 More info in the CFP: 🔗 sigtyp.github.io/ws2026-mrl.html 153
Reposted by Catherine ArnettStella Biderman @stellaathena.bsky.social · 26/05/2026"[W]hen these goods remain concentrated in the hands of a few, without adequate forms of sharing and access, a new imbalance is created that contradicts the universal destination of goods" Very cool to see the Pope endorsing @eleutherai.bsky.social's mission 2131
Catherine Arnett @catherinearnett.bsky.social · 26/03/2026I’m at #HSP2026 at MIT this week! I’ll be giving a talk Friday at 5:25pm entitled “Structural Priming Effects in Language Models are Less Human-like in Languages Other Than English”. Looking forward to chatting to everyone! 0121
Catherine Arnett @catherinearnett.bsky.social · 09/03/2026@tylerachang.bsky.social and I will be presenting the Goldfish as an oral at #LREC2026 in Mallorca! 🌴 1204
Reposted by Catherine ArnettLaurie Burchell @very-laurie.bsky.social · 25/02/2026Happening now! @pjox.bsky.social and I are giving a talk for @eleutherai.bsky.social on CommonLID, a community-driven web domain evaluation dataset for language identification. Join here: discord.gg/aYy3Se7Q?eve... Paper: arxiv.org/abs/2601.18026 @commoncrawl.bsky.socialdiscord.ggJoin the EleutherAI Discord Server!The original open science AI research collective. We started the open source LLM movement and have been pushing the boundaries of science ever since. | 33740 members 041
Reposted by Catherine Arnetteleutherai.bsky.social @eleutherai.bsky.social · 13/02/2026Announcing our latest paper: CommonLID In collaboration with @commoncrawl.bsky.social @mlcommons.org @jhu.edu we built a LID benchmark on actual Common Crawl text covering 109 languages. Existing evaluations overestimate how well LangID works on web data. arxiv.org/abs/2601.18026arxiv.orgCommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web DataLanguage identification (LID) is a fundamental step in curating multilingual corpora. However, LID models still perform poorly for many languages, especially on the noisy and heterogeneous web data of... 12212
Reposted by Catherine ArnettCommon Crawl Foundation @commoncrawl.bsky.social · 10/02/2026Language identification still proves to be a challenging task, especially for web data. In collaboration with @mlcommons.org @eleutherai.bsky.social @jhu.edu and 97 community members, we created CommonLID, a new benchmark for LangID for 100+ languages! 1105
Catherine Arnett @catherinearnett.bsky.social · 07/12/2025We will be presenting this work this afternoon! 030
Catherine Arnett @catherinearnett.bsky.social · 04/12/2025I’m presenting this today at 11am. Come find me at poster #1909! 071
Catherine Arnett @catherinearnett.bsky.social · 24/11/2025I’ll be in San Diego for #NeurIPS2025 next week! I will be presenting posters at the main conference and at the CogInterp workshop. I will also be at the Workshop on Evaluating AI in Practice at UCSD. Looking forward to chatting about multilingual NLP, evals, and tokenizers! 161
Reposted by Catherine ArnettMultilingual Representation Workshop @ EMNLP 2026 @mrl-workshop.bsky.social · 09/11/2025We have kicked off proceedings with some brief opening remarks from @catherinearnett.bsky.social 131
Reposted by Catherine ArnettEvalEval Coalition @eval-eval.bsky.social · 06/11/2025🚨 EvalEval is back - now in San Diego!🚨 🧠 Join us for the 2025 Workshop on "Evaluating AI in Practice Bridging Statistical Rigor, Sociotechnical Insights, and Ethical Boundaries" (Co-hosted with UKAISI) 📅 Dec 8, 2025 📝 Abstract due: Nov 20, 2025 Details below! ⬇️ evalevalai.com/events/works...evalevalai.com 131
Catherine Arnett @catherinearnett.bsky.social · 29/10/2025I’m so excited that Global PIQA is out! This has been a herculean effort by our 300+ contributors. The result is an extremely high-quality, culturally-specific benchmark for over 100 languages. 181
Catherine Arnett @catherinearnett.bsky.social · 28/10/2025Our #NeurIPS2025 paper shows that even comparable monolingual tokenizers have different compression rates across languages. But by getting rid of whitespace tokenization and using a custom vocab size for each language, we can reduce token premiums. Preprint out now! 1335
Reposted by Catherine ArnettWorkshop on Multilingual Data Quality Signals @wmdqs.bsky.social · 10/10/2025WMDQS is underway! Come join us in Room 520A at @colmweb.org! #COLM2025 123
Reposted by Catherine ArnettWorkshop on Multilingual Data Quality Signals @wmdqs.bsky.social · 09/10/2025In collaboration with @commoncrawl.bsky.social, MLCommons, and @eleutherai.bsky.social, the first edition of WMDQS at @colmweb.org starts tomorrow in Room 520A! We have an updated schedule on our website, including a list of all accepted papers. 133
Catherine Arnett @catherinearnett.bsky.social · 06/10/2025I’m in Montreal this week for @colmweb.org and @wmdqs.bsky.social! Looking forward to chatting about tokenizers, multilingual data, and more! #COLM2025 0120
Catherine Arnett @catherinearnett.bsky.social · 25/09/2025I have a new blog post about the so-called “tokenizer-free” approach to language modeling and why it’s not tokenizer-free at all. I also talk about why people hate tokenizers so much! 45915
Catherine Arnett @catherinearnett.bsky.social · 19/09/2025Did you know? ❌77% of language models on @hf.co are not tagged for any language 📈For 95% of languages, most models are multilingual 🚨88% of models with tags are trained on English In a new blog post, @tylerachang.bsky.social and I dig into these trends and why they matter! 👇 1132
Reposted by Catherine ArnettMultilingual Representation Workshop @ EMNLP 2026 @mrl-workshop.bsky.social · 12/09/2025We are in need of some emergency reviewers for MRL. If you are available, please fill out this form! 001
Reposted by Catherine ArnettMultilingual Representation Workshop @ EMNLP 2026 @mrl-workshop.bsky.social · 24/08/2025We extended the deadline by one day, so you have until the end of today (Aug 24) AoE to submit! Good luck! 001
Reposted by Catherine ArnettMultilingual Representation Workshop @ EMNLP 2026 @mrl-workshop.bsky.social · 18/08/2025We have over 200 volunteers now for 90+ languages! We are hoping to expand the diversity of our language coverage and are still looking for participants who speak these languages. Check out how to get involved below, and please help us spread the word! 133
Reposted by Catherine ArnettMultilingual Representation Workshop @ EMNLP 2026 @mrl-workshop.bsky.social · 05/08/2025With six weeks left before the deadline, we have had over 50 volunteers sign up to contribute for over 30 languages. If you don’t see your language represented on the map, this is your sign to get involved! 132
Catherine Arnett @catherinearnett.bsky.social · 27/07/2025I’m in Vienna all week for @aclmeeting.bsky.social and I’ll be presenting this paper on Wednesday at 11am (Poster Session 4 in HALL X4 X5)! Reach out if you want to chat about multilingual NLP, tokenizers, and open models! 0171
Reposted by Catherine ArnettPedro Ortiz Suarez @pjox.bsky.social · 21/07/2025If you want to help us improve language and cultural coverage, and build an open source LangID system, please register to our shared task on Language Identification! 💬 Registering is easy! All the details are on the shared task webpage: wmdqs.org/shared-task/ Deadline: July 23, 2025 (AoE) ⏰wmdqs.orgWMDQS: Shared Task 032
Catherine Arnett @catherinearnett.bsky.social · 19/07/2025Really grateful to the organizers for the recognition of our work! 1121
Catherine Arnett @catherinearnett.bsky.social · 10/07/2025I'll be at ICML next week for the Tokenization Workshop @tokshop.bsky.social presenting two papers: "Evaluating Morphological Alignment of Tokenizers in 70 Languages" and "BPE Stays on SCRIPT: Structured Encoding for Robust Multilingual Pretokenization". Check out the paper threads below! 1112
Catherine Arnett @catherinearnett.bsky.social · 10/07/2025MorphScore got an update! MorphScore now covers 70 languages 🌎🌍🌏 We have a new-preprint out and we will be presenting our paper at the Tokenization Workshop @tokshop.bsky.social at ICML next week! @marisahudspeth.bsky.social @brenocon.bsky.social 1134
Catherine Arnett @catherinearnett.bsky.social · 09/07/2025Just a few days left to contribute annotations before the first release of training data. We have over 17,000 document annotations so far! 131
Reposted by Catherine ArnettStella Biderman @stellaathena.bsky.social · 26/06/2025Stop by our discover server tomorrow, Friday June 27th, to hear about @catherinearnett.bsky.social's work! 272
Catherine Arnett @catherinearnett.bsky.social · 25/06/2025I'm really excited about this shared task! We hope to create a massively multilingual physical reasoning dataset in collaboration with researchers around the world 🌍 100
Catherine Arnett @catherinearnett.bsky.social · 24/06/2025The call for papers is out for the 5th edition of the Workshop on Multilingual Representation Learning which will take place in Suzhou, China co-located with EMNLP 2025! See details below! 160
Reposted by Catherine ArnettCommon Crawl Foundation @commoncrawl.bsky.social · 23/06/2025The deadline for paper submissions has been extended! The new deadline is July 3, 2025. AoE. For more information, please visit: wmdqs.orgwmdqs.org1st Workshop on Multilingual Data Quality Signals 025
Catherine Arnett @catherinearnett.bsky.social · 09/06/2025One of the biggest obstacles to improving language technologies for low-resource languages is the lack of data. To address this, we need better language identification tools. So, we're organizing a shared task on Language Identification for Web Data! #NLP #NLProc 143
Reposted by Catherine ArnettCommon Crawl Foundation @commoncrawl.bsky.social · 29/05/2025Call for papers! We are organising the 1st Workshop on Multilingual Data Quality Signals with @mlcommons.org and @eleutherai.bsky.social, held in tandem with @colmweb.org. Submit your research on multilingual data quality! Submission deadline is 23 June, more info: wmdqs.orgwmdqs.org1st Workshop on Multilingual Data Quality Signals 098
Catherine Arnett @catherinearnett.bsky.social · 06/06/2025@tylerachang.bsky.social and I are giving a talk later today at the Cambridge NLP group about the curse of multilinguality and training small models. The talk is open to the public and the link is below! 120
Catherine Arnett @catherinearnett.bsky.social · 05/06/2025My paper with @tylerachang.bsky.social and @jamichaelov.bsky.social will appear at #ACL2025NLP! The updated preprint is available on arxiv. I look forward to chatting about bilingual models in Vienna! 182
Catherine Arnett @catherinearnett.bsky.social · 03/06/2025What if we didn't use UTF-8 as a starting point for tokenization? In UTF-8, different scripts need different number of bytes. And tokenizers can create merges that lead to stranded bytes and undecodable sequences. Sander Land and I propose a novel encoding strategy that solves those problems! 1201
Catherine Arnett @catherinearnett.bsky.social · 02/06/2025If you are into tokenization and have some bandwidth this week, we need more reviewers! 001
Reposted by Catherine ArnettTokenization Workshop (TokShop) @COLM2026 @tokshop.bsky.social · 14/05/2025📣 Call for Paper Alert: TokShop @ ICML 2025 TokShop explores tokenization across all data modalities. Topics include: subword NLP techniques, multimodal approaches, multilingual challenges, post-training modification, alternative representations, and statistical perspectives.openreview.netICML 2025 Workshop TokShopWelcome to the OpenReview homepage for ICML 2025 Workshop TokShop 11712
Catherine Arnett @catherinearnett.bsky.social · 07/05/2025I’m in Paris this week to present about best practices for multilingual LLM evaluation in the open! I’m talking at PyTorch Day as part of #GOSIMParis2025. I also wrote up the content of my talk as a blog post if you’re interested - link below! 1151