Presenting at #EMNLP2025 in a moment, session on "Multilinguality and Language Diversity 2" (A301). Our paper on Tokenization Fairness: arxiv.org/abs/2509.20045
arxiv.org
Tokenization and Representation Biases in Multilingual Models on Dialectal NLP Tasks
Dialectal data are characterized by linguistic variation that appears small to humans but has a significant impact on the performance of models. This dialect gap has been related to various factors (e...