Sign in

Negar Foroutan

@negarforoutan.bsky.social
169 followers 139 following 24 posts

#NLProc PhD Student at EPFL

PostsRepliesMedia
Reposted by Negar Foroutan
Computational Linguistics @ UZH @cl-uzh.bsky.social · 10/04/2026
🔵 @negarforoutan.bsky.social, Clara Meister, Debjit Paul, @joelniklaus.bsky.social, @sinaahmadi.bsky.social, Antoine Bosselut, @ricosennrich.bsky.social . Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in Tokenization. arxiv.org/abs/2508.04796 5/7
arxiv.org
Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in Tokenization
Tokenization is the first -- and often least scrutinized -- step of most NLP pipelines. Standard algorithms for learning tokenizers rely on frequency-based objectives, which favor languages dominant in the training data and consequently leave lower-resource languages with tokenizations that are disproportionately longer, morphologically implausible, or even riddled with <UNK> placeholders. This phenomenon ultimately amplifies computational and financial inequalities between users from different language backgrounds. To remedy this, we introduce Parity-aware Byte Pair Encoding (BPE), a variant of the widely-used BPE algorithm. At every merge step, Parity-aware BPE maximizes the compression gain of the currently worst-compressed language, trading a small amount of global compression for cross-lingual parity. We find empirically that Parity-aware BPE leads to more equitable token counts across languages, with negligible impact on global compression rate and no substantial effect on language-model performance in downstream tasks.
111
Negar Foroutan @negarforoutan.bsky.social · 15/12/2025
1/ 🌍 How does mixing data from hundreds of languages affect LLM training? In our new paper "Revisiting Multilingual Data Mixtures in Language Model Pretraining" we revisit core assumptions about multilinguality using 1.1B-3B models trained on up to 400 languages. 🧵👇
196
Reposted by Negar Foroutan
Jaap Jumelet @jumelet.bsky.social · 15/10/2025
🌍Introducing BabyBabelLM: A Multilingual Benchmark of Developmentally Plausible Training Data! LLMs learn from vastly more data than humans ever experience. BabyLM challenges this paradigm by focusing on developmentally plausible data We extend this effort to 45 new languages!
14416
Reposted by Negar Foroutan
Deniz Bayazit @bayazitdeniz.bsky.social · 25/09/2025
1/🚨 New preprint How do #LLMs’ inner features change as they train? Using #crosscoders + a new causal metric, we map when features appear, strengthen, or fade across checkpoints—opening a new lens on training dynamics beyond loss curves & benchmarks. #interpretability
2156
Negar Foroutan @negarforoutan.bsky.social · 11/08/2025
🚨New Preprint! In multilingual models, the same meaning can take far more tokens in some languages, penalizing users of underrepresented languages with worse performance and higher API costs. Our Parity-aware BPE algorithm is a step toward addressing this issue: 🧵
3287
Negar Foroutan @negarforoutan.bsky.social · 23/04/2025
Stop by our poster presentation at @iclr-conf.bsky.social and discuss real multilingual evaluation! Feel free to reach out anytime during the conference! We’d love to connect!
031
Reposted by Negar Foroutan
silingao.bsky.social @silingao.bsky.social · 01/04/2025
NEW PAPER ALERT: Generating visual narratives to illustrate textual stories remains an open challenge, due to the lack of knowledge to constrain faithful and self-consistent generations. Our #CVPR2025 paper proposes a new benchmark, VinaBench, to address this challenge.
165
Reposted by Negar Foroutan
Antoine Bosselut @abosselut.bsky.social · 25/02/2025
Lots of great news out of the EPFL NLP lab these last few weeks. We'll be at @iclr-conf.bsky.social and @naaclmeeting.bsky.social in April / May to present some of our work in training dynamics, model representations, reasoning, and AI democratization. Come chat with us during the conference!
12512
Reposted by Negar Foroutan
Sepideh Mamooler@ACL🇦🇹 @smamooler.bsky.social · 17/12/2024
🚀 Introducing PICLe: a framework for in-context named-entity detection (NED) using pseudo-annotated demonstrations. 🎯 No human labeling needed—yet it outperforms few-shot learning with human annotations! #AI #NLProc #LLMs #ICL #NER
1128
Negar Foroutan @negarforoutan.bsky.social · 05/12/2024
AI is reshaping #education, but are we ready? 🚨 Our new @pnas.org article explores how #LLMs challenge traditional assessments in higher education. Instead of banning #AI, we argue for redesigning assessments to emphasize real-world problem-solving and ethical AI use.
220
Negar Foroutan @negarforoutan.bsky.social · 02/12/2024
Excited to share our work on INCLUDE! 🚀 INCLUDE sets a new standard for #LLM benchmarks—spanning 44 languages with a focus on regional knowledge and cultural context 🌍 Time for LLMs to meet the world where it is, not where it’s translated to! #Multilingual #AI #NLProc
110
Reposted by Negar Foroutan
Antoine Bosselut @abosselut.bsky.social · 26/11/2024
.@icepfl.bsky.social is hiring for multiple positions in CS (including one open call): www.epfl.ch/about/workin... Apply to come join us in Beautiful Lausanne!
epfl.ch
Open Faculty Positions
-
0129
Reposted by Negar Foroutan
Antoine Bosselut @abosselut.bsky.social · 26/11/2024
EPFL's new AI Center has a Call for applications for postdoc fellowships in all AI-related areas. Come join if you're interested in working with me and fantastic AI colleagues! Extra Perk: We actually do have lots of GPUs ! Deadline: November 29th More info at: www.epfl.ch/research/fun...
epfl.ch
EPFL AI Center Postdoctoral Fellowships
The EPFL AI Center Postdoctoral Fellowship call for proposals is now open with a deadline on 29 November 2024 (17:00 CET).Applications are encouraged from researchers at the postdoctoral level with a ...
0209