Sign in

Jaap Jumelet

@jumelet.bsky.social
762 followers 291 following 38 posts

Postdoc @rug.nl with Arianna Bisazza. Interested in NLP, interpretability, syntax, language acquisition and typology.

PostsRepliesMedia
Reposted by Jaap Jumelet
Multilingual Representation Workshop @ EMNLP 2026 @mrl-workshop.bsky.social · 20/07/2026
🎤 Speaker lineup drop! The MRL 2026 Workshop is bringing academia and industry together: @afaji.bsky.social , MBZUAI @mdlhx.bsky.social , KU Leuven @jumelet.bsky.social , Uni of Groningen Ahmet Üstün, Cohere See you all this October in Budapest! #EMNLP #MRL2026
192
Reposted by Jaap Jumelet
Francesca Padovani @frap98.bsky.social · 14/04/2026
I’m very happy to share that my latest paper on the 𝐚𝐜𝐪𝐮𝐢𝐬𝐢𝐭𝐢𝐨𝐧 𝐨𝐟 𝐯𝐞𝐫𝐛 𝐦𝐞𝐚𝐧𝐢𝐧𝐠, tested on models trained under CDL data vs ADL data , has been accepted for an 𝐨𝐫𝐚𝐥 𝐩𝐫𝐞𝐬𝐞𝐧𝐭𝐚𝐭𝐢𝐨𝐧 at the upcoming edition of 𝐂𝐨𝐠𝐒𝐜𝐢, which will take place at the end of July in Rio de Janeiro.💃
1246
Reposted by Jaap Jumelet
Vitalii Hirak @v-hirak.bsky.social · 08/02/2026
Our paper has been accepted to #EACL2026 main conference! Together with @jumelet.bsky.social and @arianna-bis.bsky.social, we study the effect of target language typology on the difficulty of state-of-the-art neural machine translation. arXiv preprint: arxiv.org/abs/2602.03551 1/6 Our findings ⬇️
arxiv.org
Assessing the Impact of Typological Features on Multilingual Machine Translation in the Age of Large Language Models
Despite major advances in multilingual modeling, large quality disparities persist across languages. Besides the obvious impact of uneven training resources, typological properties have also been prop...
1151
Reposted by Jaap Jumelet
GroNLP @gronlp.bsky.social · 18/12/2025
👀 Look what 🎅 has broght just before Christmas 🎁: a brand new Research Master in Natural Language Processing at @facultyofartsug.bsky.social @rug.nl Program: www.rug.nl/masters/natu... Applications (2026/2027) are open! Come and study with us (you will also learn why we have a 🐮 in our logo)
rug.nl
Natural Language Processing
How do you build Large Language Models? How do humans experience Natural Language Processing (NLP) applications in their daily lives? And how can we...
02515
Reposted by Jaap Jumelet
Leonie Weissweiler @weissweiler.bsky.social · 11/12/2025
🧑‍🔬I’m recruiting PhD students in Natural Language Processing @unileipzig.bsky.social Computer Science, together with @scadsai.bsky.social! Topics include, but aren’t limited to: 🔎Linguistic Interpretability 🌍Multilingual Evaluation 📖Computational Typology Please share! #NLProc #NLP
14225
Reposted by Jaap Jumelet
Leonie Weissweiler @weissweiler.bsky.social · 19/11/2025
📢Out now in NEJLT!📢 In each of these sentences, a verb that doesn't usually encode motion is being used to convey that an object is moving to a destination. Given that these usages are rare, complex, and creative, we ask: Do LLMs understand what's going on in them? 🧵1/7
2153
Reposted by Jaap Jumelet
Jennifer Hu @jennhu.bsky.social · 10/11/2025
New work to appear @ TACL! Language models (LMs) are remarkably good at generating novel well-formed sentences, leading to claims that they have mastered grammar. Yet they often assign higher probability to ungrammatical strings than to grammatical strings. How can both things be true? 🧵👇
Screenshot of a figure with two panels, labeled (a) and (b). The caption reads: "Figure 1: (a) Illustration of messages (left) and strings (right) in toy domain. Blue = grammatical strings. Red = ungrammatical strings. (b) Surprisal (negative log probability) assigned to toy strings by GPT-2."
29220
Jaap Jumelet @jumelet.bsky.social · 06/11/2025
I'm in Suzhou to present our work on MultiBLiMP, Friday @ 11:45 in the Multilinguality session (A301)! Come check it out if your interested in multilingual linguistic evaluation of LLMs (there will be parse trees on the slides! There's still use for syntactic structure!) arxiv.org/abs/2504.02768
0277
Reposted by Jaap Jumelet
GroNLP @gronlp.bsky.social · 27/10/2025
With only a week left for #EMNLP2025, we are happy to announce all the works we 🐮 will present 🥳 - come and say "hi" to our posters and presentations during the Main and the co-located events (*SEM and workshops) See you in Suzhou ✈️
accepted papers at main conference and findingsaccepted papers at TACL and workshops
0165
Jaap Jumelet @jumelet.bsky.social · 15/10/2025
For more information check out the website, paper, and datasets: Website: babylm.github.io/babybabellm/ Paper: arxiv.org/pdf/2510.10159 We hope BabyBabelLM will continue as a 'living resource', fostering both more efficient NLP methods, and opening ways for cross-lingual computational linguistics!
babylm.github.io
BabyBabelLM
010
Jaap Jumelet @jumelet.bsky.social · 15/10/2025
Next to our training resources, we also release an evaluation pipeline that assess different aspects of language learning. We present results for various simple baseline models, but hope this can serve as a starting point for a multilingual BabyLM challenge in future years!
100
Jaap Jumelet @jumelet.bsky.social · 15/10/2025
To deal with data imbalances, we divide languages into three Tiers. This better enables cross-lingual studies and makes it possible for low-resource languages to be a part of BabyBabelLM as well.
100
Jaap Jumelet @jumelet.bsky.social · 15/10/2025
With a fantastic team of international collaborators we have developed a pipeline for creating LM training data from resources that children are exposed to. We release this pipeline and welcome new contributions! Website: babylm.github.io/babybabellm/ Paper: arxiv.org/pdf/2510.10159
111
Jaap Jumelet @jumelet.bsky.social · 15/10/2025
🌍Introducing BabyBabelLM: A Multilingual Benchmark of Developmentally Plausible Training Data! LLMs learn from vastly more data than humans ever experience. BabyLM challenges this paradigm by focusing on developmentally plausible data We extend this effort to 45 new languages!
14416
Jaap Jumelet @jumelet.bsky.social · 01/09/2025
Wij speelden als kind (in Breda) vaak "1 keer tets", waar je een voetbal maximaal 1 keer mocht laten stuiteren; ik had ook geen idee dat dat een Brabants woord was.
010
Jaap Jumelet @jumelet.bsky.social · 01/08/2025
Happening now at the SIGTYP poster session! Come talk to Leonie and me about MultiBLiMP!
1202
Reposted by Jaap Jumelet
Jelle Zuidema 🟥 @wzuidema.bsky.social · 29/07/2025
I'll be in Vienna only from tomorrow, but today my star PhD student Marianne is already presenting some of our work: BLIMP-NL, in which we create a large new dataset for syntactic evaluation of Dutch LLMs, and learn a lot about dataset creation, LLM evaluation and grammatical abilities on the way.
1111
Jaap Jumelet @jumelet.bsky.social · 02/07/2025
Congrats and good luck in Canada!
010
Reposted by Jaap Jumelet
Arianna Bisazza @arianna-bis.bsky.social · 19/06/2025
Proud to introduce TurBLiMP, the 1st benchmark of minimal pairs for free-order, morphologically rich Turkish language! Pre-print: arxiv.org/abs/2506.13487 Fruit of an almost year-long project by amazing MS student @ezgibasar.bsky.social in collab w/ @frap98.bsky.social and @jumelet.bsky.social
arxiv.org
TurBLiMP: A Turkish Benchmark of Linguistic Minimal Pairs
We introduce TurBLiMP, the first Turkish benchmark of linguistic minimal pairs, designed to evaluate the linguistic abilities of monolingual and multilingual language models (LMs). Covering 16 linguis...
1112
Reposted by Jaap Jumelet
Casper Albers 🟥 @casperalbers.nl · 13/06/2025
Ik snap niet dat hier niet meer ophef over is: Het binnenhalen van Amerikaanse wetenschappers wordt betaalt door Nederlandse academici geen inflatiecorrectie op hun salaris te geven. 1/2
56535
Jaap Jumelet @jumelet.bsky.social · 12/06/2025
Ohh cool! Nice to see the interactions-as-structure idea I had back in 2021 is still being explored!
030
Reposted by Jaap Jumelet
Catherine Arnett @catherinearnett.bsky.social · 05/06/2025
My paper with @tylerachang.bsky.social and @jamichaelov.bsky.social will appear at #ACL2025NLP! The updated preprint is available on arxiv. I look forward to chatting about bilingual models in Vienna!
182
Reposted by Jaap Jumelet
Francesca Padovani @frap98.bsky.social · 30/05/2025
“Child-Directed Language Does Not Consistently Boost Syntax Learning in Language Models” I’m happy to share that the preprint of my first PhD project is now online! 🎊 Paper: arxiv.org/abs/2505.23689
arxiv.org
Child-Directed Language Does Not Consistently Boost Syntax Learning in Language Models
Seminal work by Huebner et al. (2021) showed that language models (LMs) trained on English Child-Directed Language (CDL) can reach similar syntactic abilities as LMs trained on much larger amounts of ...
26117
Reposted by Jaap Jumelet
Neil Renic @ncrenic.bsky.social · 28/05/2025
"A well-delivered lecture isn’t primarily a delivery system for information. It is an ignition point for curiosity, all the better for being experienced in an audience." Marvellous defence of the increasingly maligned university experience by @patporter76.bsky.social thecritic.co.uk/university-a...
thecritic.co.uk
University: a good idea | Patrick Porter | The Critic Magazine
A former student of mine has penned an attack on universities, derived from their own disappointing experience studying Politics and International Relations at the place where I ply my trade. In short...
05819
Reposted by Jaap Jumelet
Miryam de Lhoneux @mdlhx.bsky.social · 16/05/2025
Interested in multilingual tokenization in #NLP? Lisa Beinborn and I are hiring! PhD candidate position in Göttingen, Germany: www.uni-goettingen.de/de/644546.ht... PostDoc position in Leuven, Belgium: www.kuleuven.be/personeel/jo... Deadline 6th of June
uni-goettingen.de
Stellen OBP - Georg-August-Universität Göttingen
Webseiten der Georg-August-Universität Göttingen
22513
Reposted by Jaap Jumelet
BlackboxNLP @blackboxnlp.bsky.social · 15/05/2025
BlackboxNLP, the leading workshop on interpretability and analysis of language models, will be co-located with EMNLP 2025 in Suzhou this November! 📆 This edition will feature a new shared task on circuits/causal variable localization in LMs, details here: blackboxnlp.github.io/2025/task
3218
Reposted by Jaap Jumelet
Leshem (Legend) Choshen @EMNLP @lchoshen.bsky.social · 09/05/2025
Close your books, test time! The evaluation pipelines are out, baselines are released & the challenge is on There is still time to join and We are excited to learn from you on pretraining and human-model gaps *Don't forget to fastEval on checkpoints github.com/babylm/evalu... 📈🤖🧠 #AI #LLMS
0104
Reposted by Jaap Jumelet
Seth Aycock @sethjsa.bsky.social · 25/04/2025
Pleased to announce our paper was accepted at ICLR 2025 as a Spotlight! I will present our poster on Saturday April 26, 3-5pm, Poster #241. Hope to see you there! arxiv.org/abs/2409.19151
arxiv.org
Can LLMs Really Learn to Translate a Low-Resource Language from One Grammar Book?
Extremely low-resource (XLR) languages lack substantial corpora for training NLP models, motivating the use of all available resources such as dictionaries and grammar books. Machine Translation from ...
0173
Jaap Jumelet @jumelet.bsky.social · 23/04/2025
Scherp geschreven en geheel mee eens, maar beetje wrang wel dat de boodschap zich achter een paywall van 450 euro bevindt :') (dank voor de screenshots!)
110
Reposted by Jaap Jumelet
Jirui Qi @jiruiqi.bsky.social · 11/04/2025
✨ New Paper ✨ [1/] Retrieving passages from many languages can boost retrieval augmented generation (RAG) performance, but how good are LLMs at dealing with multilingual contexts in the prompt? 📄 Check it out: arxiv.org/abs/2504.00597 (w/ @arianna-bis.bsky.social @Raquel_Fernández) #NLProc
145
Jaap Jumelet @jumelet.bsky.social · 17/04/2025
That is definitely possible indeed, and a potential confounding factor. In RuBLiMP, a Russian benchmark, they defined a way to validate this based on LM probs, but we left that open for future work. The poor performance on low-res langs shows they're definitely not trained on all of UD though!
110
Reposted by Jaap Jumelet
Jaap Jumelet @jumelet.bsky.social · 07/04/2025
✨New paper ✨ Introducing 🌍MultiBLiMP 1.0: A Massively Multilingual Benchmark of Minimal Pairs for Subject-Verb Agreement, covering 101 languages! We present over 125,000 minimal pairs and evaluate 17 LLMs, finding that support is still lacking for many languages. 🧵⬇️
37622
Reposted by Jaap Jumelet
Arianna Bisazza @arianna-bis.bsky.social · 08/04/2025
Modern LLMs "speak" hundreds of languages... but do they really? Multilinguality claims are often based on downstream tasks like QA & MT, while *formal* linguistic competence remains hard to gauge in lots of languages Meet MultiBLiMP! (joint work w/ @jumelet.bsky.social & @weissweiler.bsky.social)
2216
Jaap Jumelet @jumelet.bsky.social · 07/04/2025
Joint work with @weissweiler.bsky.social and @arianna-bis.bsky.social. Check out the full paper at arxiv.org/abs/2504.02768! We have released all our data on huggingface huggingface.co/datasets/jum... We hope to extend this pipeline to many more phenomena!
arxiv.org
MultiBLiMP 1.0: A Massively Multilingual Benchmark of Linguistic Minimal Pairs
We introduce MultiBLiMP 1.0, a massively multilingual benchmark of linguistic minimal pairs, covering 101 languages, 6 linguistic phenomena and containing more than 125,000 minimal pairs. Our minimal ...
061
Jaap Jumelet @jumelet.bsky.social · 07/04/2025
Person agreement is easier to model than Gender or Number. Sentences with higher overall perplexity lead to less accurate judgements, and models are more likely to pick the wrong inflection if it is split into more tokens. Surprisingly, subject-verb distance has no effect.
130
Jaap Jumelet @jumelet.bsky.social · 07/04/2025
We find that boosting specific languages works, but only if you pre-, and not post-train: EuroLLM outperforms same size Llama3 on its target languages, but Aya is not significantly better. Neither of them outperform Llama3 significantly on a language not intentionally included.
120
Jaap Jumelet @jumelet.bsky.social · 07/04/2025
We evaluate 17 Language Models, among them Llama 3, Aya, and Gemma 3. Overall, Llama3 70B and Gemma 27B perform best, but the monolingual 500M Goldfish models significantly outperform them in 14 languages! Base models consistently outperform their instruction-tuned counterparts.
250
Jaap Jumelet @jumelet.bsky.social · 07/04/2025
We create 125,000 pairs for 101 languages and six types of agreement, resulting in high diversity across phenomena, typological families, geography, amount of resources available, sentence length, and word frequencies. 43 of our languages are not Indo-European.
110
Jaap Jumelet @jumelet.bsky.social · 07/04/2025
MultiBLiMP is created automatically using Universal Dependencies and Universal Morphology. We search for subject-verb or -participle pairs with our target features Number, Person, and Gender in UD, then insert the word with the opposite feature value to form a minimal pair.
120
Jaap Jumelet @jumelet.bsky.social · 07/04/2025
✨New paper ✨ Introducing 🌍MultiBLiMP 1.0: A Massively Multilingual Benchmark of Minimal Pairs for Subject-Verb Agreement, covering 101 languages! We present over 125,000 minimal pairs and evaluate 17 LLMs, finding that support is still lacking for many languages. 🧵⬇️
37622
Jaap Jumelet @jumelet.bsky.social · 01/04/2025
Quite some papers in this direction recently, really fruitful direction (and shameless self-plug, did this in 2021 already): aclanthology.org/2021.finding... aclanthology.org/2024.tacl-1.... arxiv.org/abs/2407.04593 aclanthology.org/2024.emnlp-m...
aclanthology.org
Filtered Corpus Training (FiCT) Shows that Language Models Can Generalize from Indirect Evidence
Abhinav Patil, Jaap Jumelet, Yu Ying Chiu, Andy Lapastora, Peter Shen, Lexie Wang, Clevis Willrich, Shane Steinert-Threlkeld. Transactions of the Association for Computational Linguistics, Volume 12. ...
072
Jaap Jumelet @jumelet.bsky.social · 31/03/2025
Fantastic paper!! Fascinating findings, really cool to see this whole corpus modifying setup being so useful for investigating these questions.
130
Reposted by Jaap Jumelet
Qing Yao @qyao.bsky.social · 31/03/2025
LMs learn argument-based preferences for dative constructions (preferring recipient first when it’s shorter), consistent with humans. Is this from memorizing preferences in training? New paper w/ @kanishka.bsky.social , @weissweiler.bsky.social , @kmahowald.bsky.social arxiv.org/abs/2503.20850
examples from direct and prepositional object datives with short-first and long-first word orders: 
DO (long first): She gave the boy who signed up for class and was excited it.
PO (short first): She gave it to the boy who signed up for class and was excited.
DO (short first): She gave him the book that everyone was excited to read.
PO (long-first): She gave the book that everyone was excited to read to him.
1177
Jaap Jumelet @jumelet.bsky.social · 13/03/2025
To make my original assumption a bit more explicit: I expected the meta-linguistic ability to be something that arises mostly from post-training, and not from a larger amount of pre-training data / larger model, whereas your result seems to suggest the latter.
120
Reposted by Jaap Jumelet
Siyuan Song @siyuansong.bsky.social · 12/03/2025
New preprint w/ @jennhu.bsky.social @kmahowald.bsky.social : Can LLMs introspect about their knowledge of language? Across models and domains, we did not find evidence that LLMs have privileged access to their own predictions. 🧵(1/8)
26116
Jaap Jumelet @jumelet.bsky.social · 13/03/2025
This is a fascinating result! I would have expected there to be a more noticeable difference between pre- and post-trained models. Did you observe any other differences in meta-linguistic performance for the post-trained model wrt to its base model?
110
Jaap Jumelet @jumelet.bsky.social · 10/03/2025
Would love to join that!
010
Reposted by Jaap Jumelet
Arianna Bisazza @arianna-bis.bsky.social · 01/03/2025
Excited to be traveling to Estonia for the 1st time to give a keynote @nodalida.bsky.social. I'll talk about using NNs to study language evolution & acquisition. A teaser: It won't be about LLMs 🙃 Also I've just moved from X, so this was my very first post... Pls help out by connecting with me!
4599
Reposted by Jaap Jumelet
Leonie Weissweiler @weissweiler.bsky.social · 20/02/2025
✨New paper✨ Linguistic evaluations of LLMs often implicitly assume that language is generated by symbolic rules. In a new position paper, @adelegoldberg.bsky.social, @kmahowald.bsky.social and I argue that languages are not Lego sets, and evaluations should reflect this! arxiv.org/pdf/2502.13195
16819
Jaap Jumelet @jumelet.bsky.social · 20/02/2025
That's not my point. You don't judge the performance of a system based on the behaviour of an outdated version.
0100