Sign in

Catherine Arnett

@catherinearnett.bsky.social
4.1K followers 624 following 127 posts

NLP Researcher at EleutherAI, PhD UC San Diego Linguistics. Working on making language technologies more equitable. 📍Oxford, UK. She/her. catherinearnett.github.io

PostsRepliesMedia
Reposted by Catherine Arnett
Marco @mcognetta.bsky.social · 30/09/2026
🚨 [Token][ization] Paper Alert 🚨 Tokenization is a wildly understudied area of language modeling despite it having effects across all of NLP. Over the past ~8 months, 32 (!) tokenizer researchers put together the most comprehensive survey of the field. Check it out!
112935
Reposted by Catherine Arnett
Leshem (Legend) Choshen @EMNLP @lchoshen.bsky.social · 28/08/2026
Will the same model become a different player depending on the interface language? We report our findings in our paper "Skill Issue: Are Skills Language-Invariant in LLMs?". And with it, TextArena is now multilingual 🌍 with 65 games in 193 languages. 🧵
1172
Catherine Arnett @catherinearnett.bsky.social · 28/08/2026
Xiulin has been doing some extremely interesting and important work about how to go about actually comparable language model evaluation across models and across languages. We have some recommendations - check them out in our new preprint and catch the paper at #EMNLP2026!
1110
Catherine Arnett @catherinearnett.bsky.social · 20/08/2026
As a personal update, I have moved to the city of Oxford where I’ll continue working from EleutherAI! Looking forward to meeting people in the area!
1150
Reposted by Catherine Arnett
Multilingual Representation Workshop @ EMNLP 2026 @mrl-workshop.bsky.social · 28/07/2026
We are super impressed by the submissions for Phase 1 of the MRL Shared Task! :) In order for us to thoroughly review the datasets, we will extend the deadline to August 14! The audio collection phase will now happen in September. Please stay tuned here (or on Discord) for more updates!
011
Reposted by Catherine Arnett
James Michaelov @jamichaelov.bsky.social · 11/06/2026
Seems like a good time to share our new preprint about model openness! (with @catherinearnett.bsky.social @tylerachang.bsky.social Pamela D. Rivière, Samuel M. Taylor @camrobjones.bsky.social @seantrott.bsky.social @rplevy.bsky.social Ben Bergen, and Micah Altman): arxiv.org/abs/2603.26539
1193
Reposted by Catherine Arnett
Stella Biderman @stellaathena.bsky.social · 10/06/2026
In film, "we'll fix it in post" is what you say when something went wrong on set and you don't want to redo it. AI research has made it our entire methodology: train the model, then patch whatever comes out. Our new ICML oral argues this can't be the basis of a science of AI. 🧵
310823
Catherine Arnett @catherinearnett.bsky.social · 08/06/2026
This is a great opportunity for students and early-stage researchers. Please contribute if you can!
020
Catherine Arnett @catherinearnett.bsky.social · 02/06/2026
The new and expanded version of Global PIQA is out with over twice as many items. Well done to all the contributors!
1113
Reposted by Catherine Arnett
Multilingual Representation Workshop @ EMNLP 2026 @mrl-workshop.bsky.social · 27/05/2026
📢 Call for Papers: 6th Multilingual Representation Learning Workshop at EMNLP in Budapest, Hungary! Join us and submit your works relating to multilingual NLP Speakers to be announced, so stay tuned! 👀 More info in the CFP: 🔗 sigtyp.github.io/ws2026-mrl.html
153
Reposted by Catherine Arnett
Stella Biderman @stellaathena.bsky.social · 26/05/2026
"[W]hen these goods remain concentrated in the hands of a few, without adequate forms of sharing and access, a new imbalance is created that contradicts the universal destination of goods" Very cool to see the Pope endorsing @eleutherai.bsky.social's mission
2131
Catherine Arnett @catherinearnett.bsky.social · 26/03/2026
I’m at #HSP2026 at MIT this week! I’ll be giving a talk Friday at 5:25pm entitled “Structural Priming Effects in Language Models are Less Human-like in Languages Other Than English”. Looking forward to chatting to everyone!
0121
Catherine Arnett @catherinearnett.bsky.social · 09/03/2026
@tylerachang.bsky.social and I will be presenting the Goldfish as an oral at #LREC2026 in Mallorca! 🌴
1204
Reposted by Catherine Arnett
Laurie Burchell @very-laurie.bsky.social · 25/02/2026
Happening now! @pjox.bsky.social and I are giving a talk for @eleutherai.bsky.social on CommonLID, a community-driven web domain evaluation dataset for language identification. Join here: discord.gg/aYy3Se7Q?eve... Paper: arxiv.org/abs/2601.18026 @commoncrawl.bsky.social
discord.gg
Join the EleutherAI Discord Server!
The original open science AI research collective. We started the open source LLM movement and have been pushing the boundaries of science ever since. | 33740 members
041
Reposted by Catherine Arnett
eleutherai.bsky.social @eleutherai.bsky.social · 13/02/2026
Announcing our latest paper: CommonLID In collaboration with @commoncrawl.bsky.social @mlcommons.org @jhu.edu we built a LID benchmark on actual Common Crawl text covering 109 languages. Existing evaluations overestimate how well LangID works on web data. arxiv.org/abs/2601.18026
arxiv.org
CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data
Language identification (LID) is a fundamental step in curating multilingual corpora. However, LID models still perform poorly for many languages, especially on the noisy and heterogeneous web data of...
12212
Reposted by Catherine Arnett
Common Crawl Foundation @commoncrawl.bsky.social · 10/02/2026
Language identification still proves to be a challenging task, especially for web data. In collaboration with @mlcommons.org @eleutherai.bsky.social @jhu.edu and 97 community members, we created CommonLID, a new benchmark for LangID for 100+ languages!
Examples of mislabeled web text by existing LangID systems. A full text version is available on the blog post below.Examples of mislabeled web text by existing LangID systems. A full text version is available on the blog post below.
1105
Catherine Arnett @catherinearnett.bsky.social · 07/12/2025
We will be presenting this work this afternoon!
030
Catherine Arnett @catherinearnett.bsky.social · 04/12/2025
I’m presenting this today at 11am. Come find me at poster #1909!
071
Catherine Arnett @catherinearnett.bsky.social · 24/11/2025
I’ll be in San Diego for #NeurIPS2025 next week! I will be presenting posters at the main conference and at the CogInterp workshop. I will also be at the Workshop on Evaluating AI in Practice at UCSD. Looking forward to chatting about multilingual NLP, evals, and tokenizers!
161
Reposted by Catherine Arnett
Multilingual Representation Workshop @ EMNLP 2026 @mrl-workshop.bsky.social · 09/11/2025
We have kicked off proceedings with some brief opening remarks from @catherinearnett.bsky.social
131
Reposted by Catherine Arnett
EvalEval Coalition @eval-eval.bsky.social · 06/11/2025
🚨 EvalEval is back - now in San Diego!🚨 🧠 Join us for the 2025 Workshop on "Evaluating AI in Practice Bridging Statistical Rigor, Sociotechnical Insights, and Ethical Boundaries" (Co-hosted with UKAISI) 📅 Dec 8, 2025 📝 Abstract due: Nov 20, 2025 Details below! ⬇️ evalevalai.com/events/works...
evalevalai.com
131
Catherine Arnett @catherinearnett.bsky.social · 29/10/2025
I’m so excited that Global PIQA is out! This has been a herculean effort by our 300+ contributors. The result is an extremely high-quality, culturally-specific benchmark for over 100 languages.
181
Catherine Arnett @catherinearnett.bsky.social · 28/10/2025
Our #NeurIPS2025 paper shows that even comparable monolingual tokenizers have different compression rates across languages. But by getting rid of whitespace tokenization and using a custom vocab size for each language, we can reduce token premiums. Preprint out now!
1335
Reposted by Catherine Arnett
Workshop on Multilingual Data Quality Signals @wmdqs.bsky.social · 10/10/2025
WMDQS is underway! Come join us in Room 520A at @colmweb.org! #COLM2025
123
Reposted by Catherine Arnett
Workshop on Multilingual Data Quality Signals @wmdqs.bsky.social · 09/10/2025
In collaboration with @commoncrawl.bsky.social, MLCommons, and @eleutherai.bsky.social, the first edition of WMDQS at @colmweb.org starts tomorrow in Room 520A! We have an updated schedule on our website, including a list of all accepted papers.
133
Catherine Arnett @catherinearnett.bsky.social · 06/10/2025
I’m in Montreal this week for @colmweb.org and @wmdqs.bsky.social! Looking forward to chatting about tokenizers, multilingual data, and more! #COLM2025
Name tag with “Anti Anti Tokenizer Club” pin on lanyard
0120
Catherine Arnett @catherinearnett.bsky.social · 25/09/2025
I have a new blog post about the so-called “tokenizer-free” approach to language modeling and why it’s not tokenizer-free at all. I also talk about why people hate tokenizers so much!
45915
Catherine Arnett @catherinearnett.bsky.social · 19/09/2025
Did you know? ❌77% of language models on @hf.co are not tagged for any language 📈For 95% of languages, most models are multilingual 🚨88% of models with tags are trained on English In a new blog post, @tylerachang.bsky.social and I dig into these trends and why they matter! 👇
1132
Reposted by Catherine Arnett
Multilingual Representation Workshop @ EMNLP 2026 @mrl-workshop.bsky.social · 12/09/2025
We are in need of some emergency reviewers for MRL. If you are available, please fill out this form!
001
Reposted by Catherine Arnett
Multilingual Representation Workshop @ EMNLP 2026 @mrl-workshop.bsky.social · 24/08/2025
We extended the deadline by one day, so you have until the end of today (Aug 24) AoE to submit! Good luck!
001
Reposted by Catherine Arnett
Multilingual Representation Workshop @ EMNLP 2026 @mrl-workshop.bsky.social · 18/08/2025
We have over 200 volunteers now for 90+ languages! We are hoping to expand the diversity of our language coverage and are still looking for participants who speak these languages. Check out how to get involved below, and please help us spread the word!
We are still actively looking for volunteers speaking the following languages (or other languages not listed):
Afrikaans, Aymara, Basque, Bosnian, Breton, Burmese, Cebuano, Guarani, Haitian Creole, Hmong, Hungarian, Icelandic, Inuktitut, Irish, Karakalpak, Khmer, Kirghiz, Lao, Latvian, Macedonian, Malagasy, Maltese, Maori, Mongolian, Nahuatl, Navajo/Diné, Norwegian Nynorsk, Quechua, Romanian, Samoan, Scottish Gaelic, Shona, Somali, Tatar, Tibetan, Tigrinya, Waray, Walloon, Welsh, Yiddish, Zulu.
133
Reposted by Catherine Arnett
Multilingual Representation Workshop @ EMNLP 2026 @mrl-workshop.bsky.social · 05/08/2025
With six weeks left before the deadline, we have had over 50 volunteers sign up to contribute for over 30 languages. If you don’t see your language represented on the map, this is your sign to get involved!
132
Catherine Arnett @catherinearnett.bsky.social · 27/07/2025
I’m in Vienna all week for @aclmeeting.bsky.social and I’ll be presenting this paper on Wednesday at 11am (Poster Session 4 in HALL X4 X5)! Reach out if you want to chat about multilingual NLP, tokenizers, and open models!
0171
Reposted by Catherine Arnett
Pedro Ortiz Suarez @pjox.bsky.social · 21/07/2025
If you want to help us improve language and cultural coverage, and build an open source LangID system, please register to our shared task on Language Identification! 💬 Registering is easy! All the details are on the shared task webpage: wmdqs.org/shared-task/ Deadline: July 23, 2025 (AoE) ⏰
wmdqs.org
WMDQS: Shared Task
032
Catherine Arnett @catherinearnett.bsky.social · 19/07/2025
Really grateful to the organizers for the recognition of our work!
1121
Catherine Arnett @catherinearnett.bsky.social · 10/07/2025
I'll be at ICML next week for the Tokenization Workshop @tokshop.bsky.social presenting two papers: "Evaluating Morphological Alignment of Tokenizers in 70 Languages" and "BPE Stays on SCRIPT: Structured Encoding for Robust Multilingual Pretokenization". Check out the paper threads below!
1112
Catherine Arnett @catherinearnett.bsky.social · 10/07/2025
MorphScore got an update! MorphScore now covers 70 languages 🌎🌍🌏 We have a new-preprint out and we will be presenting our paper at the Tokenization Workshop @tokshop.bsky.social at ICML next week! @marisahudspeth.bsky.social @brenocon.bsky.social
1134
Catherine Arnett @catherinearnett.bsky.social · 09/07/2025
Just a few days left to contribute annotations before the first release of training data. We have over 17,000 document annotations so far!
131
Reposted by Catherine Arnett
Stella Biderman @stellaathena.bsky.social · 26/06/2025
Stop by our discover server tomorrow, Friday June 27th, to hear about @catherinearnett.bsky.social's work!
272
Catherine Arnett @catherinearnett.bsky.social · 25/06/2025
I'm really excited about this shared task! We hope to create a massively multilingual physical reasoning dataset in collaboration with researchers around the world 🌍
100
Catherine Arnett @catherinearnett.bsky.social · 24/06/2025
The call for papers is out for the 5th edition of the Workshop on Multilingual Representation Learning which will take place in Suzhou, China co-located with EMNLP 2025! See details below!
160
Reposted by Catherine Arnett
Common Crawl Foundation @commoncrawl.bsky.social · 23/06/2025
The deadline for paper submissions has been extended! The new deadline is July 3, 2025. AoE. For more information, please visit: wmdqs.org
wmdqs.org
1st Workshop on Multilingual Data Quality Signals
025
Catherine Arnett @catherinearnett.bsky.social · 09/06/2025
One of the biggest obstacles to improving language technologies for low-resource languages is the lack of data. To address this, we need better language identification tools. So, we're organizing a shared task on Language Identification for Web Data! #NLP #NLProc
143
Reposted by Catherine Arnett
Common Crawl Foundation @commoncrawl.bsky.social · 29/05/2025
Call for papers! We are organising the 1st Workshop on Multilingual Data Quality Signals with @mlcommons.org and @eleutherai.bsky.social, held in tandem with @colmweb.org. Submit your research on multilingual data quality! Submission deadline is 23 June, more info: wmdqs.org
wmdqs.org
1st Workshop on Multilingual Data Quality Signals
098
Catherine Arnett @catherinearnett.bsky.social · 06/06/2025
@tylerachang.bsky.social and I are giving a talk later today at the Cambridge NLP group about the curse of multilinguality and training small models. The talk is open to the public and the link is below!
120
Catherine Arnett @catherinearnett.bsky.social · 05/06/2025
My paper with @tylerachang.bsky.social and @jamichaelov.bsky.social will appear at #ACL2025NLP! The updated preprint is available on arxiv. I look forward to chatting about bilingual models in Vienna!
182
Catherine Arnett @catherinearnett.bsky.social · 03/06/2025
What if we didn't use UTF-8 as a starting point for tokenization? In UTF-8, different scripts need different number of bytes. And tokenizers can create merges that lead to stranded bytes and undecodable sequences. Sander Land and I propose a novel encoding strategy that solves those problems!
1201
Catherine Arnett @catherinearnett.bsky.social · 02/06/2025
If you are into tokenization and have some bandwidth this week, we need more reviewers!
001
Reposted by Catherine Arnett
Tokenization Workshop (TokShop) @COLM2026 @tokshop.bsky.social · 14/05/2025
📣 Call for Paper Alert: TokShop @ ICML 2025 TokShop explores tokenization across all data modalities. Topics include: subword NLP techniques, multimodal approaches, multilingual challenges, post-training modification, alternative representations, and statistical perspectives.
openreview.net
ICML 2025 Workshop TokShop
Welcome to the OpenReview homepage for ICML 2025 Workshop TokShop
11712
Catherine Arnett @catherinearnett.bsky.social · 07/05/2025
I’m in Paris this week to present about best practices for multilingual LLM evaluation in the open! I’m talking at PyTorch Day as part of #GOSIMParis2025. I also wrote up the content of my talk as a blog post if you’re interested - link below!
1151