Reposted by Tyler ChangMultilingual Representation Workshop @ EMNLP 2026 @mrl-workshop.bsky.social · 02/06/2026We are releasing an expanded version of Global PIQA! It now covers 141 language varieties and includes parallel and non-parallel splits. We are also releasing an updated preprint. 152
Tyler Chang @tylerachang.bsky.social · 29/10/2025Very very excited that Global PIQA is out! This was an incredible effort by 300+ researchers from 65 countries. The resulting dataset is a high-quality, participatory, and culturally-specific benchmark for over 100 languages. 040
Reposted by Tyler ChangCatherine Arnett @catherinearnett.bsky.social · 19/09/2025Did you know? ❌77% of language models on @hf.co are not tagged for any language 📈For 95% of languages, most models are multilingual 🚨88% of models with tags are trained on English In a new blog post, @tylerachang.bsky.social and I dig into these trends and why they matter! 👇 1132
Reposted by Tyler ChangMultilingual Representation Workshop @ EMNLP 2026 @mrl-workshop.bsky.social · 18/08/2025We have over 200 volunteers now for 90+ languages! We are hoping to expand the diversity of our language coverage and are still looking for participants who speak these languages. Check out how to get involved below, and please help us spread the word! 133
Reposted by Tyler ChangMultilingual Representation Workshop @ EMNLP 2026 @mrl-workshop.bsky.social · 05/08/2025With six weeks left before the deadline, we have had over 50 volunteers sign up to contribute for over 30 languages. If you don’t see your language represented on the map, this is your sign to get involved! 132
Tyler Chang @tylerachang.bsky.social · 25/06/2025We're organizing a shared task to develop a multilingual physical commonsense reasoning evaluation dataset! Details on how to submit are at: sigtyp.github.io/st2025-mrl.h...sigtyp.github.ioMRL 2025 Shared Task on Multilingual Physical Reasoning Datasets 040
Tyler Chang @tylerachang.bsky.social · 25/04/2025Presenting our work on training data attribution for pretraining this morning: iclr.cc/virtual/2025... -- come stop by in Hall 2/3 #526 if you're here at ICLR!iclr.ccICLR Poster Scalable Influence and Fact Tracing for Large Language Model PretrainingICLR 2025 140
Tyler Chang @tylerachang.bsky.social · 13/12/2024We scaled training data attribution (TDA) methods ~1000x to find influential pretraining examples for thousands of queries in an 8B-parameter LLM over the entire 160B-token C4 corpus! medium.com/people-ai-re... 2357
Reposted by Tyler ChangCatherine Arnett @catherinearnett.bsky.social · 22/11/2024The Goldfish models were trained on byte-premium-scaled dataset sizes, such that if a language needs more bytes to encode a given amount of information, we scaled up the dataset according the byte premium. Read about how we (@tylerachang.bsky.social) trained the models: arxiv.org/pdf/2408.10441 151
Reposted by Tyler ChangCatherine Arnett @catherinearnett.bsky.social · 15/11/2024Tyler Chang and my paper got awarded outstanding paper at #EMNLP2024! Thanks to the award committee for the recognition! 1321