Reposted by Gianluca VicoMarco @mcognetta.bsky.social · 30/09/2026🚨 [Token][ization] Paper Alert 🚨 Tokenization is a wildly understudied area of language modeling despite it having effects across all of NLP. Over the past ~8 months, 32 (!) tokenizer researchers put together the most comprehensive survey of the field. Check it out! 112433
Reposted by Gianluca VicoTokenization Workshop (TokShop) @COLM2026 @tokshop.bsky.social · 14/05/2026TokShop will be at #COLM2026! 🗓️ October 9th, 2026 📍 San Francisco, USA More details and a call for papers coming soon. 0118
Reposted by Gianluca VicoInstitute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 30/03/2026#EACL2026 is over. Thank for representing @ufal.mff.cuni.cz and congrats on the work you have presented! @abyste.bsky.social, @gianlucavico.bsky.social, @patuchen.bsky.social, @jlibovicky.bsky.social, @zdenekkasner.cz, @tuetschek.bsky.social, @namlh201.bsky.social, @jjon19.bsky.social et al. 1101
Reposted by Gianluca VicoInstitute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 27/03/2026On the weekend, #EACL2026 continues with workshops and @ufal.mff.cuni.cz folks present their research 👇 Also, don't miss @tuetschek.bsky.social's keynote talk on How (Not) to Find Errors in LLM Outputs at the LowResMT workshop in the morning. 182
Reposted by Gianluca VicoInstitute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 29/03/2026Crowdsourcing Piedmontese to Test LLMs on Non-Standard Orthography by @gianlucavico.bsky.social and @jlibovicky.bsky.social aclanthology.org/2026.vardial... New Piedmontese dataset tests tokenization, classification & translation! 🗣️ 141
Reposted by Gianluca VicoInstitute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 24/03/2026#EACL2026 in Rabat 🇲🇦 starts tomorrow and @ufal.mff.cuni.cz folks will present their research. Don't miss our presentations 👇 1132
Reposted by Gianluca VicoIlker Kesen @ilkerkesen.bsky.social · 17/03/2026📢I'm organizing a BoF session at #EACL2026 called Tokenization & Beyond, aiming to gather researchers exploring tokenization and alternatives such as byte-level and pixel-based approaches. Sign up using the form if you're interested! #NLProc @eaclmeeting.bsky.social 1119
Gianluca Vico @gianlucavico.bsky.social · 10/11/2025We’re collecting crowd-sourced translations in Piedmontese and Neapolitan. 🎯 Goal: see how well LLMs understand these languages. 👉 Participate here (in IT🇮🇹): - Piedmontese: quest.ms.mff.cuni.cz/crowd-transl... - Neapolitan: quest.ms.mff.cuni.cz/crowd-transl... Anyone can join, no need to be fluent!quest.ms.mff.cuni.czWelcome to CrowdTranslation 052
Reposted by Gianluca VicoJindřich Libovický @jlibovicky.bsky.social · 25/08/2025🧵 We're releasing CUS-QA - a new benchmark for testing LLMs on regional knowledge! Find out what your model knows about Czechia 🇨🇿, Slovakia 🇸🇰, and Ukraine 🇺🇦! 👉 Textual and visual questions, answers, and human judgment on model outputs! huggingface.co/datasets/ufa... www.arxiv.org/abs/2507.22752huggingface.coufal/cus-qa · Datasets at Hugging FaceWe’re on a journey to advance and democratize artificial intelligence through open source and open science. 1163
Reposted by Gianluca VicoInstitute of Formal and Applied Linguistics @ufal.mff.cuni.cz · 26/08/2025Our team of 10 is at #MTMarathon2025 in Helsinki 🇫🇮, a week-long meeting of machine translation researchers, developers. ✅ Posters presented ✅ Now working on cool collaborative projects with researchers from around the world. #MachineTranslation #NLP 0144