Reposted by Antoine BosselutMete @mismayil.bsky.social · 13/04/2026LLMs can retrieve knowledge — but can they connect it in *creative* ways to solve problems? Introducing CresOWLve 🦉, a new benchmark that evaluates creative problem-solving over real-world knowledge, using puzzles that require multiple creative thinking strategies.👇 132
Reposted by Antoine BosselutNegar Foroutan @negarforoutan.bsky.social · 15/12/20251/ 🌍 How does mixing data from hundreds of languages affect LLM training? In our new paper "Revisiting Multilingual Data Mixtures in Language Model Pretraining" we revisit core assumptions about multilinguality using 1.1B-3B models trained on up to 400 languages. 🧵👇 196
Reposted by Antoine BosselutUKP Lab @ukplab.bsky.social · 19/12/2025🎤 Prof. Iryna Gurevych, Distinguished Professor at the National Center for Cybersecurity @athenecenter.bsky.social and Director of the UKP Lab at @tuda.bsky.social, delivered a keynote on “How to Make AI-Native Internet Content Secure? Coping with Synthetic and Misleading Data.” 141
Reposted by Antoine BosselutEPFL School of Computer and Communication Sciences @icepfl.bsky.social · 11/11/2025🎉 Congratulations to Assistant Professors @abosselut.bsky.social (IC), @bunnech.bsky.social (IC & SV), and @mschrimpf.bsky.social (IC & SV) for being selected as #AI2050 Early Career Fellows by @schmidtsciences.bsky.social ! 🔗 Full article: actu.epfl.ch/news/epfl-pr...actu.epfl.ch 1102
Reposted by Antoine Bosselutchenhaotan.bsky.social @chenhaotan.bsky.social · 03/11/2025Recruiting PhDs & postdocs for: 🤖 agents "taking over" science (hypogenic.ai and 📌) 🧪 Real scientists ➡️AI (e.g., materials, chem, physics) 📜 Theory + incentives for H-AI collab & credit (e.g., formalizing tacit knowledge) new adventures for me, 🔄 if you can! 🙌 chenhaot.com/recruiting.h...chenhaot.comChenhao Tan's Homepage - recruitingChenhao Tan's Homepage 093
Antoine Bosselut @abosselut.bsky.social · 14/10/2025If you're interested in doing a postdoc at @icepfl.bsky.social , there's still time to apply for the @epfl-ai-center.bsky.social postdoctoral fellowships. Apart from this, I'm also recruiting postdocs in developing novel training algorithms for reasoning models and agentic AI. 182
Antoine Bosselut @abosselut.bsky.social · 10/10/2025Join us again at #MELT workshop (520D) at #COLM2025 to hear from @ImanolSchlag about #Apertus, the largest multilingual LLM trained on over 1000 languages. 020
Antoine Bosselut @abosselut.bsky.social · 10/10/2025Kicking off #MELT workshop at #COLM2025 with Monojit Choudhury talking about "Meta-Cultural Competence: What LLMs Should Know About Culture to Serve the Next Billion Users" ! 050
Antoine Bosselut @abosselut.bsky.social · 10/10/2025Come join us in 520D (all the way down the hall and around the corner) at #COLM2025 for the first workshop on multilingual and equitable language technologies! 021
Reposted by Antoine BosselutTiago Pimentel @tpimentel.bsky.social · 01/10/2025Very happy this paper got accepted to NeurIPS 2025 as a Spotlight! 😁 Main takeaway: In mechanistic interpretability, we need assumptions about how DNNs encode concepts in their representations (eg, the linear representation hypothesis). Without them, we can claim any DNN implements any algorithm! 0254
Reposted by Antoine BosselutAaron Mueller @amuuueller.bsky.social · 01/10/2025What's the right unit of analysis for understanding LLM internals? We explore in our mech interp survey (a major update from our 2024 ms). We’ve added more recent work and more immediately actionable directions for future work. Now published in Computational Linguistics! 24115
Reposted by Antoine BosselutDeniz Bayazit @bayazitdeniz.bsky.social · 25/09/20251/🚨 New preprint How do #LLMs’ inner features change as they train? Using #crosscoders + a new causal metric, we map when features appear, strengthen, or fade across checkpoints—opening a new lens on training dynamics beyond loss curves & benchmarks. #interpretability 2156
Reposted by Antoine BosselutMete @mismayil.bsky.social · 22/09/2025💡Can we optimize LLMs to be more creative? Introducing Creative Preference Optimization (CrPO) and MuCE (Multi-task Creativity Evaluation Dataset). Result: More novel, diverse, surprising text—without losing quality! 📝 Appearing at #EMNLP2025 164
Antoine Bosselut @abosselut.bsky.social · 03/09/2025The next generation of open LLMs should be inclusive, compliant, and multilingual by design. That’s why we @icepfl.bsky.social @ethz.ch @cscsch.bsky.social ) built Apertus. 2248
Reposted by Antoine BosselutEPFL AI Center @epfl-ai-center.bsky.social · 02/09/2025EPFL, @ethz.ch and the @cscsch.bsky.social released Apertus today, Switzerland’s first large-scale, open, multilingual language model — a milestone in generative AI for transparency and diversity. Find out more here: ai.epfl.ch/apertus-a-fu... @abosselut.bsky.social @icepfl.bsky.socialai.epfl.chApertus: a fully open, transparent, multilingual language model - EPFL AI CenterEPFL, ETH Zurich and the Swiss National Supercomputing Centre (CSCS) released Apertus today, Switzerland’s first large-scale, open, multilingual language model — a milestone in generative AI for trans... 0197
Reposted by Antoine BosselutEPFL School of Computer and Communication Sciences @icepfl.bsky.social · 02/09/2025EPFL, ETH Zurich & CSCS just released Apertus, Switzerland’s first fully open-source large language model. Trained on 15T tokens in 1,000+ languages, it’s built for transparency, responsibility & the public good. Read more: actu.epfl.ch/news/apertus... 15429
Reposted by Antoine BosselutAlexander Doria @dorialexander.bsky.social · 02/09/2025Very happy to see that Pleias multilingual data processing pipelines have contributed to the largest open pretraining project in Europe. From their tech report: huggingface.co/swiss-ai/Ape... 23010
Reposted by Antoine BosselutReto Vogt @rvgt.ch · 02/09/2025Die Schweiz steigt ins Rennen der grossen Sprachmodelle ein. Unter dem Namen #Apertus veröffentlichen @ethz.ch, @icepfl.bsky.social und das @cscsch.bsky.social das erste vollständig offene, mehrsprachige #LLM des Landes. Fürs MAZ habe ich Apertus kurz analysiert: www.maz.ch/news/apertus...maz.chApertus: ein neues Sprachmodell für die Schweiz 3277
Reposted by Antoine Bosselutkyunghyuncho.bsky.social @kyunghyuncho.bsky.social · 12/08/2025recently gave a talk on <Reality Checks> at two venues, and discussed (and rambled) about how leaderboard chasing is awesome (and we want it to continue) but that this isn't easy because everyone (me! me! me!) wants to write more papers. the link to the slide deck in the reply. 3245
Reposted by Antoine BosselutNegar Foroutan @negarforoutan.bsky.social · 11/08/2025🚨New Preprint! In multilingual models, the same meaning can take far more tokens in some languages, penalizing users of underrepresented languages with worse performance and higher API costs. Our Parity-aware BPE algorithm is a step toward addressing this issue: 🧵 3287
Antoine Bosselut @abosselut.bsky.social · 04/08/2025The EPFL NLP lab is looking to hire a postdoctoral researcher on the topic of designing, training, and evaluating multilingual LLMs: docs.google.com/document/d/1... Come join our dynamic group in beautiful Lausanne!docs.google.comEPFL NLP Postdoctoral Scholar Posting - Swiss AI LLMsThe EPFL Natural Language Processing (NLP) lab is looking to hire a postdoctoral researcher candidate in the area of multilingual LLM design, training, and evaluation. This postdoctoral position is as... 02012
Reposted by Antoine BosselutAbhilasha Ravichander @lasha.bsky.social · 22/07/2025📣 Life update: Thrilled to announce that I’ll be starting as faculty at the Max Planck Institute for Software Systems this Fall! I’ll be recruiting PhD students in the upcoming cycle, as well as research interns throughout the year: lasharavichander.github.io/contact.html 139212
Reposted by Antoine BosselutEPFL AI Center @epfl-ai-center.bsky.social · 09/07/2025EPFL and ETH Zürich are building together a Swiss made LLM from scratch. Fully open and multilingual, the model is trained on CSCS's supercomputer "Alps" and supports sovereign, transparent, and responsible AI in Switzerland and beyond. Read more here: ai.epfl.ch/a-language-m... #ResponsibleAIai.epfl.chA language model built for the public good - EPFL AI CenterETH Zurich and EPFL will release a large language model (LLM) developed on public infrastructure. Trained on the “Alps” supercomputer at the Swiss National Supercomputing Centre (CSCS), the new LLM ma... 0103
Antoine Bosselut @abosselut.bsky.social · 23/06/2025Check out Silin's paper done in collaboration with Apple on reinforcing abstract thinking in reasoning traces! 020
Antoine Bosselut @abosselut.bsky.social · 18/06/2025Check out @bkhmsi.bsky.social 's great work on mixture-of-expert models that are specialized to represent the behavior of known brain networks. 031
Reposted by Antoine BosselutEPFL AI Center @epfl-ai-center.bsky.social · 02/06/2025Many AI models speak dozens of languages, but do they grasp cultural context? 🗣️🌍 The INCLUDE benchmark from EPFL's NLP Lab and @cohereforai.bsky.social reveal that there is still a gap... 👉 Find out how benchmarks like INCLUDE can help make AI truly inclusive: actu.epfl.ch/news/beyond-...actu.epfl.chBeyond translation – making AI multiculturalA team of international researchers led by EPFL developed a multilingual benchmark to determine Large Language Models ability to grasp cultural context. 041
Reposted by Antoine BosselutVered Shwartz @veredshwartz.bsky.social · 27/05/2025I guess that now that I have 1% of my Twitter followers follow me here 😅, I should announce it here too for those of you no longer checking Twitter: my nonfiction book, "Lost in Automatic Translation" is coming out this July: lostinautomatictranslation.com. I'm very excited to share it with you! 17615
Reposted by Antoine BosselutDebjit Paul @debjit-paul.bsky.social · 01/05/2025Super excited to share that our paper "A Logical Fallacy-Informed Framework for Argument Generation" has received the Outstanding Paper Award 🎉🎉 at NAACL 2025! Paper: aclanthology.org/2025.naacl-l... Code: github.com/lucamouchel/... #NAACL2025aclanthology.orgA Logical Fallacy-Informed Framework for Argument GenerationLuca Mouchel, Debjit Paul, Shaobo Cui, Robert West, Antoine Bosselut, Boi Faltings. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Lingu... 042
Reposted by Antoine BosselutBadr AlKhamissi @bkhmsi.bsky.social · 30/04/2025Excited to be at #NAACL2025 in Albuquerque! I’ll be presenting our paper “The LLM Language Network” as an Oral tomorrow at 2:00 PM in Ballroom C, hope to see you there! Looking forward to all the discussions! 🎤 🧠 1104
Reposted by Antoine BosselutMartin Jaggi @mjaggi.bsky.social · 23/04/2025Using the 'right' data can hugely speed up LLM training, but how to find the best training data in the vast sea of a whole web crawl? We propose a simple classifier-based selection, enabling multilingual LLMs 🧵 182
Reposted by Antoine BosselutAngelika Romanou @agromanou.bsky.social · 23/04/2025If you’re at @iclr-conf.bsky.social this week, come check out our spotlight poster INCLUDE during the Thursday 3:00–5:30pm session! I will be there to chat about all things multilingual & multicultural evaluation. Feel free to reach out anytime during the conference. I’d love to connect! 042
Reposted by Antoine BosselutPrithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 14/04/2025The Wordplay Workshop is back! 5th edition with EMNLP in Suzhou this Dec. We're also hosting a competition this time on making more realistic LLM powered NPCs in games! As always come by and chat all things text agents! wordplay-workshop.github.io www.aicrowd.com/challenges/c... 1172
Reposted by Antoine Bosselutsilingao.bsky.social @silingao.bsky.social · 01/04/2025NEW PAPER ALERT: Generating visual narratives to illustrate textual stories remains an open challenge, due to the lack of knowledge to constrain faithful and self-consistent generations. Our #CVPR2025 paper proposes a new benchmark, VinaBench, to address this challenge. 165
Reposted by Antoine BosselutConference on Language Modeling @colmweb.org · 20/03/2025A bit of a mess around the conflict of COLM with the ARR (and to lesser degree ICML) reviews release. We feel this is creating a lot of pressure and uncertainty. So, we are pushing our deadlines: Abstracts due March 22 AoE (+48hr) Full papers due March 28 AoE (+24hr) Plz RT 🙏 33631
Reposted by Antoine BosselutAkhil Arora @akhilarora.bsky.social · 18/03/2025I am recruiting 2 PhD students for Fall'25 @csaudk.bsky.social to work on bleeding-edge topics in #NLProc #LLMs #AIAgents (e.g. LLM reasoning, knowledge-seeking agents, and more). Details: www.cs.au.dk/~clan/openings Deadline: May 1, 2025 Please boost! cc: @aicentre.dk @wikiresearch.bsky.socialcs.au.dkOpen positions and projects### Open semester and Master's projects If you're an AU student looking for a semester project, a Bachelor project, or an MS thesis project, please refer to [this list](projects). ### Prospective PhD ... 03022
Reposted by Antoine BosselutConference on Language Modeling @colmweb.org · 09/03/2025COLM's @juliakreutzer.bsky.social and @abosselut.bsky.social will hold two paper submission Q&A sessions. We run a simple process, but figured this can help authors, especially first-time authors. March 12: dateful.com/eventlink/14... March 13: dateful.com/eventlink/83... Plz RT 🙏 053
Reposted by Antoine BosselutBadr AlKhamissi @bkhmsi.bsky.social · 05/03/2025🚨 New Preprint!! LLMs trained on next-word prediction (NWP) show high alignment with brain recordings. But what drives this alignment—linguistic structure or world knowledge? And how does this alignment evolve during training? Our new paper explores these questions. 👇🧵 15923
Antoine Bosselut @abosselut.bsky.social · 05/03/2025We have released the March 2025 Swiss AI Call for Large Projects! Call Document: docs.google.com/document/d/1... Highlights: - 1-year projects - 2 M CHF for personnel - 10 M GPUh Let's change the scale of academic research! Deadline: March 31st, 2025docs.google.comSwiss AI - Call for Large Projects March 2025Swiss AI Initiative - Call for Large Grants Overview AI advances are progressing at a speed and scale never seen before, with unprecedented opportunities for disruptive breakthrough applications. Howe... 082
Antoine Bosselut @abosselut.bsky.social · 25/02/2025Lots of great news out of the EPFL NLP lab these last few weeks. We'll be at @iclr-conf.bsky.social and @naaclmeeting.bsky.social in April / May to present some of our work in training dynamics, model representations, reasoning, and AI democratization. Come chat with us during the conference! 12512
Reposted by Antoine BosselutMete @mismayil.bsky.social · 20/02/2025Check out our paper [arxiv.org/abs/2410.12656] for more details. Huge shoutouts to my amazing collaborators @defnecirci.bsky.social, @jonnesaleva.bsky.social, Hale Sirin, Abdullatif Koksal, Bhuwan Dhingra, @abosselut.bsky.social, Duygu Ataman, Lonneke van der Plas. 021
Reposted by Antoine BosselutYoav Artzi @yoavartzi.com · 18/02/2025We recently pushed an update to our in-context RL paper. Usually, updates don't justify a post, but this one is exceptionally contentful -> 🧵 tl;dr: all the findings are stronger, and the behaviors are super cool! arxiv.org/abs/2410.05362 1184
Reposted by Antoine BosselutJekaterina Novikova @j-novikova-nlp.bsky.social · 23/01/2025Our paper is accepted to ICLR! INCLUDE: Evaluating Multilingual LLMs with Regional Knowledge (arxiv.org/abs/2411.19799) A benchmark of ~200k QA pairs across 44 languages, capturing real-world cultural nuances. A collaborative effort led by @cohereforai.bsky.social, with contributors worldwide. /1 1114
Reposted by Antoine BosselutYoav Artzi @yoavartzi.com · 14/02/2025Increasingly frustrated with the misunderstanding of basic experimental design in the community. People confuse experimental testbeds designed to show effect and answer specific research questions vs. the applicability of the approach. Leading to experimental demands that stifle research 2261
Antoine Bosselut @abosselut.bsky.social · 06/01/2025For all the gripes I often hear about #NLP ARR, I don't think it's mentioned enough how great it is to have submission deadlines set in stone. We used to just get a CFP 4-6 months before the deadline describing deadline, review period, rebuttal period (as many other confs still do). 1150
Reposted by Antoine BosselutBadr AlKhamissi @bkhmsi.bsky.social · 19/12/2024🚨 New Paper! Can neuroscience localizers uncover brain-like functional specializations in LLMs? 🧠🤖 Yes! We analyzed 18 LLMs and found units mirroring the brain's language, theory of mind, and multiple demand networks! w/ @gretatuckute.bsky.social, @abosselut.bsky.social, @mschrimpf.bsky.social 🧵👇 210325
Reposted by Antoine BosselutSepideh Mamooler@ACL🇦🇹 @smamooler.bsky.social · 17/12/2024🚀 Introducing PICLe: a framework for in-context named-entity detection (NED) using pseudo-annotated demonstrations. 🎯 No human labeling needed—yet it outperforms few-shot learning with human annotations! #AI #NLProc #LLMs #ICL #NER 1128
Reposted by Antoine BosselutRobert Csordas @robertcsordas.bsky.social · 12/12/2024Come visit our poster "MoEUT: Mixture-of-Experts Universal Transformers" on Friday at 4:30 in East Exhibit Hall A-C #1907 on #NeurIPS2024. With Kazuki Irie, Jürgen Schmidhuber, Christopher Potts and @chrmanning.bsky.social. 1145
Reposted by Antoine BosselutSara Hooker @sarahooker.bsky.social · 05/12/2024Is MMLU Western-centric? 🤔 As part of a massive cross-institutional collaboration: 🗽Find MMLU is heavily overfit to western culture 🔍 Professional annotation of cultural sensitivity data 🌍 Release improved Global-MMLU 42 languages 📜 Paper: arxiv.org/pdf/2412.03304 📂 Data: hf.co/datasets/Coh... 75912
Reposted by Antoine BosselutGuilherme Penedo @guilherme.hf.co · 08/12/2024Announcing 🥂 FineWeb2: A sparkling update with 1000s of 🗣️languages. We applied the same data-driven approach that led to SOTA English performance in🍷 FineWeb to thousands of languages. 🥂 FineWeb2 has 8TB of compressed text data and outperforms other datasets. 17619