Sign in

Paul Lerner

@lernerp.bsky.social
50 followers 56 following 50 posts

Research Scientist @ NuMind paullerner.github.io

PostsRepliesMedia
Paul Lerner @lernerp.bsky.social · 28/08/2026
The second twist is that you don’t even have to provide data: a meta-task of autollm is to detect relevant documents from large corpora (e.g. CommonCrawl). (this twist is yet to be implemented :))
000
Paul Lerner @lernerp.bsky.social · 28/08/2026
autollm works like any AutoML libraries you would expect: input data, you get a trained model. The twist is that you don’t have to provide annotations along with the data, it is annotated automatically by an LLM (e.g. ChatGPT).
100
Paul Lerner @lernerp.bsky.social · 28/08/2026
I'm starting this new personal project: autollm, a python library for easy data distillation of LLMs for text classification, information extraction, and open-ended tasks github.com/PaulLerner/a...
github.com
120
Paul Lerner @lernerp.bsky.social · 04/06/2026
Hi folks, I'm searching for a Research Scientist position in Paris starting from September! Let me know if you hear about an opportunity :)
000
Paul Lerner @lernerp.bsky.social · 14/05/2026
Great discussion at LREC :)
000
Reposted by Paul Lerner
Inria Paris NLP (ALMAnaCH team) @inriaparisnlp.bsky.social · 08/05/2026
Parallel Corpora of Scholarly Documents for English-French Machine Translation Ziqian Peng, Lichao Zhu, Rachel Bawden, Maud Bénard, Éric de la Clergerie, Mathilde Huguin, Natalie Kübler, Paul Lerner, Alexandra Mestivier & François Yvon 📅 11th May | 16:30–16:54 | BUCC (remote)
101
Reposted by Paul Lerner
MLIA ISIR @mlia-isir.bsky.social · 06/05/2026
➡️ Can Multimodal LLMs Generate Pedagogical Questions? Thomas Gerald, Sahar Ghannay, Julie Lascar, Paul Lerner @lernerp.bsky.social, Anne Vilnat in collaboration with LISN @lisnlab.bsky.social x.com/LISNLAB 📅 Thurs., 14 May, 11:00 - 12:40 (long, poster)
021
Reposted by Paul Lerner
MLIA ISIR @mlia-isir.bsky.social · 06/05/2026
➡️ Assessing the Political Fairness of Multilingual LLMs: A Case Study based on a 21-way Multiparallel EuroParl Dataset Paul Lerner @lernerp.bsky.social, François Yvon @yvofr.bsky.social 📅 Wed., 13 May, 11:40 (long, oral) 📖 arxiv.org/abs/2510.20508
021
Reposted by Paul Lerner
MLIA ISIR @mlia-isir.bsky.social · 06/05/2026
➡️ Parallel Corpora of Scholarly Documents for English-French Machine Translation #BUCC Z. Peng, @lichaozhu.bsky.social @rachelbawden.bsky.social @maudbenard.bsky.social Éric de la Clergerie, @mathildehuguin.bsky.social @nataliekubler.bsky.social @lernerp.bsky.social A. Mestivier & @yvofr.bsky.social
122
Paul Lerner @lernerp.bsky.social · 05/05/2026
Happy to chat if you're at LREC next week, I'll be presenting Wednesday at 11:40 in Session O4 "Evaluation, Validation, Quality Assurance and Benchmarking Methodologies"
000
Reposted by Paul Lerner
Inria Paris NLP (ALMAnaCH team) @inriaparisnlp.bsky.social · 03/04/2026
We are very happy to announce our next seminar: Paul Lerner @lernerp.bsky.social (ISIR, Sorbonne Université & CNRS) "Controlling Linguistic Variability in Large Language Models" on Friday 10th April 2026, 11am CET. Details here 👉 almanach.inria.fr/seminars-en....
ALMAnaCH seminar: Paul Lerner, “Controlling Linguistic Variability in Large Language Models”, 10/04/2026
132
Reposted by Paul Lerner
Leonie Weissweiler @weissweiler.bsky.social · 11/12/2025
🧑‍🔬I’m recruiting PhD students in Natural Language Processing @unileipzig.bsky.social Computer Science, together with @scadsai.bsky.social! Topics include, but aren’t limited to: 🔎Linguistic Interpretability 🌍Multilingual Evaluation 📖Computational Typology Please share! #NLProc #NLP
14225
Reposted by Paul Lerner
MLIA ISIR @mlia-isir.bsky.social · 24/11/2025
The team meeting of the week was presented by Alexandre Vérine, from PSL, about "Quality and Diversity in generative models through the lens of f-divergences." Thanks a lot for this interesting talk!
001
Reposted by Paul Lerner
MLIA ISIR @mlia-isir.bsky.social · 28/10/2025
Accepted to a Workshop (1/2): "Self-Retrieval from Distant Contexts for Document-Level Machine Translation", accepted to the Conference on Machine Translation (WMT25), from @ziqianpeng.bsky.social, @rachelbawden.bsky.social, @yvofr.bsky.social
102
Paul Lerner @lernerp.bsky.social · 06/11/2025
There's many directions where this could go, multilingual, low-resource language, interpretability, depending on your profile, and the internship may lead to a PhD, provided we get funding!
011
Paul Lerner @lernerp.bsky.social · 06/11/2025
As we found in aclanthology.org/2025.coling-... that BPE-based LLMs (i.e. pretty much all LLMs) did not handle prefixations well
aclanthology.org
Unlike “Likely”, “Unlike” is Unlikely: BPE-based Segmentation hurts Morphological Derivations in LLMs
Paul Lerner, François Yvon. Proceedings of the 31st International Conference on Computational Linguistics. 2025.
111
Paul Lerner @lernerp.bsky.social · 06/11/2025
Basically the idea is to extend www.pnas.org/doi/10.1073/... to see how well LLMs model competition between affixes, not only suffixes (e.g. -ity vs. -ness) but also prefixes (e.g. un- vs. non-)
pnas.org
Derivational morphology reveals analogical generalization in large language models | PNAS
What mechanisms underlie linguistic generalization in large language models (LLMs)? This question has attracted considerable attention, with most s...
121
Paul Lerner @lernerp.bsky.social · 06/11/2025
Come work with @yvofr.bsky.social @weissweiler.bsky.social and me at @mlia-isir.bsky.social for a M2 internship on Assessing the Morphological Competence of LLMs! For 5-6 months from February or March 2026. Paid 600€/month
132
Paul Lerner @lernerp.bsky.social · 24/10/2025
What's the plural of "LLM-as-a-Judge"?
000
Paul Lerner @lernerp.bsky.social · 24/10/2025
now on arxiv arxiv.org/abs/2510.20508
arxiv.org
Assessing the Political Fairness of Multilingual LLMs: A Case Study based on a 21-way Multiparallel EuroParl Dataset
The political biases of Large Language Models (LLMs) are usually assessed by simulating their answers to English surveys. In this work, we propose an alternative framing of political biases, relying o...
000
Paul Lerner @lernerp.bsky.social · 23/10/2025
work done with @yvofr.bsky.social as part of the Democratic Commons programme, many thanks to our colleagues at Make, Sciences Po, and Sorbonne! about.make.org/democratic-c...
about.make.org
Landing Page
100
Paul Lerner @lernerp.bsky.social · 23/10/2025
the dataset and code are available github.com/PaulLerner/2...
github.com
GitHub - PaulLerner/21-EuroParl: Dataset and code for the paper "Assessing the Political Fairness of Multilingual LLMs: A Case Study based on a 21-way Multiparallel EuroParl Dataset" (Lerner and Yvon,...
Dataset and code for the paper "Assessing the Political Fairness of Multilingual LLMs: A Case Study based on a 21-way Multiparallel EuroParl Dataset" (Lerner and Yvon, 2025) - PaulLerner/...
100
Paul Lerner @lernerp.bsky.social · 23/10/2025
here's what one example of the dataset looks like, there are 72,234 just like this one (I regret my multimodal days where there were pictures in my papers)
100
Paul Lerner @lernerp.bsky.social · 23/10/2025
We find that LLMs translate some political parties unfairly using a new version of EuroParl, fully multi-parallel and including (political) metadata hal.science/hal-05328251
hal.science
Assessing the Political Fairness of Multilingual LLMs: A Case Study based on a 21-way Multiparallel EuroParl Dataset
The political biases of Large Language Models (LLMs) are usually assessed by simulating their answers to English surveys. In this work, we propose an alternative framing of political biases, relying on principles of fairness in multilingual translation. We systematically compare the translation quality of speeches in the European Parliament (EP), observing systematic differences with majority parties from left, center, and right being better translated than outsider parties. This study is made possible by a new, 21-way multiparallel version of EuroParl, the parliamentary proceedings of the EP, which includes the political affiliations of each speaker. The dataset consists of 1.5M sentences for a total of 40M words and 249M characters. It covers three years, 1000+ speakers, 7 countries, 12 EU parties, 25 EU committees, and hundreds of national parties.
110
Paul Lerner @lernerp.bsky.social · 15/10/2025
I tried for a pythonic library, have a look at the example notebook colab.research.google.com/github/PaulL...
colab.research.google.com
Google Colab
000
Paul Lerner @lernerp.bsky.social · 15/10/2025
🤔 ppllm is benchmarked against: - a vllm-based implementation: 4.15 times faster! - a naive hugging face implementation, which does not sort texts by length: 4.61 times faster!
100
Paul Lerner @lernerp.bsky.social · 15/10/2025
🤔 ppllm implements windowed PPL, which allows to compute the PPL of arbitrarily long texts. It aims to be feature complete for many information-theoretic metrics, including Perplexity (PPL), Surprisal, and bits per character (BPC), and their word-level counterparts.
100
Paul Lerner @lernerp.bsky.social · 15/10/2025
introducing 🤔 ppllm, a Python Library to Compute LLM's Perplexity and Surprisal github.com/PaulLerner/p...
github.com
GitHub - PaulLerner/ppllm: 🤔 A Python Library to Compute LLM's Perplexity and Surprisal
🤔 A Python Library to Compute LLM's Perplexity and Surprisal - PaulLerner/ppllm
100
Paul Lerner @lernerp.bsky.social · 04/09/2025
make.org/FR/consultat...
make.org
Contribute to current consultations - Comment l’IA peut-elle améliorer la vie des Français en limitant les risques ? - Make.org
Finding proposals is easier when working together. Discover a democratic place where you can discuss the big issues you care about, submit your proposals concerning them and vote on proposals proposed...
000
Reposted by Paul Lerner
Pasquale Minervini @neuralnoise.com · 05/07/2025
"in 2025 we will have flying cars" 😂😂😂
839991
Paul Lerner @lernerp.bsky.social · 07/07/2025
Work done with Laurène Cave, @haldaume3.bsky.social, Léo Labat, Gaël Lejeune, Pierre-Antoine Lequeu, @bpiwowar.bsky.social, Nazanin Shafiabadi and @yvofr.bsky.social, read the paper here talnarchives.atala.org/ateliers/202... Any feedback is appreciated :)
talnarchives.atala.org
000
Paul Lerner @lernerp.bsky.social · 07/07/2025
Last week, I presented my work on "Assessing the Political Biases of Multilingual LLMs" at the EALM workshop @ TALN 2025 ! Thanks again to the ANR Diké project for organizing the workshop
100
Reposted by Paul Lerner
MLIA ISIR @mlia-isir.bsky.social · 10/06/2025
📢 🎉 The team has one paper accepted to #MTsummit2025! "Investigating Length Issues in Document-level Machine Translation" by @ziqianpeng.bsky.social, @rachelbawden.bsky.social and @yvofr.bsky.social in collaboration with @inriaparisnlp.bsky.social 📍 Geneva | 🗓️ 23-27,June 📕 arxiv.org/abs/2412.17592
arxiv.org
Investigating Length Issues in Document-level Machine Translation
Transformer architectures are increasingly effective at processing and generating very long chunks of texts, opening new perspectives for document-level machine translation (MT). In this work, we chal...
111
Paul Lerner @lernerp.bsky.social · 16/06/2025
"meticulously" is so absent from this list (from aclanthology.org/2025.coling-... )
000
Paul Lerner @lernerp.bsky.social · 16/06/2025
Am I the only reviewer that actually fills this "Reviewer Checklist"? And why do Area Chairs never answer when the paper needs to be desk-rejected? And reviews are due in 3 days 🫠
000
Reposted by Paul Lerner
MLIA ISIR @mlia-isir.bsky.social · 10/06/2025
For the EALM Workshop "On Assessing the Political Biases of Multilingual Large Language Models" by @lernerp.bsky.social Laurène Cave, @haldaume3.bsky.social Léo Labat, Gaël Lejeune, Pierre-Antoine Lequeu, @bpiwowar.bsky.social Nazanin Shafiabadi and yvofr.bsky.social, collaborated with the STIH lab
102
Paul Lerner @lernerp.bsky.social · 07/05/2025
@mdlhx.bsky.social PS: I just found out we can actually share projects outside of CNRS plmlatex.math.cnrs.fr/6632958859wh...
plmlatex.math.cnrs.fr
PLMlatex, Éditeur LaTeX en ligne
Un éditeur LaTeX en ligne facile à utiliser. Pas d’installation, collaboration en temps réel, gestion des versions, des centaines de modèles de documents LaTeX, et plus encore.
010
Paul Lerner @lernerp.bsky.social · 06/05/2025
CNRS provides plmlatex.math.cnrs.fr that covers most of the features. I guess it's not so complicated to host (the software is open source)
plmlatex.math.cnrs.fr
Identifiant
Un éditeur LaTeX en ligne facile à utiliser. Pas d’installation, collaboration en temps réel, gestion des versions, des centaines de modèles de documents LaTeX, et plus encore.
110
Paul Lerner @lernerp.bsky.social · 28/03/2025
", I am" 🤔 Don't you think this would increase the imbalance in multilingual LLMs?
000
Paul Lerner @lernerp.bsky.social · 20/02/2025
Amazed at what a COLING paper could look like in the 80's
010
Paul Lerner @lernerp.bsky.social · 11/02/2025
Hope you enjoyed our poster at #AISummit! I'm standing next to Pierre-Antoine Lequeu, @salimhafid.bsky.social, and @manonberriche.bsky.social but there's more people involved! Zoom-in to read their names or learn more about the project here about.make.org/democratic-c...
031
Paul Lerner @lernerp.bsky.social · 25/01/2025
accurate quote for NLP researchers visiting Louvre Abu Dhabi after COLING 2025
010
Paul Lerner @lernerp.bsky.social · 22/01/2025
Makes me think about this quiz I made with former PhD students at LISN, can you spot the fake BERT-based models?
000
Paul Lerner @lernerp.bsky.social · 22/01/2025
btw "DagoBERT" is my all time favorite pun with "BERT" names
100
Paul Lerner @lernerp.bsky.social · 22/01/2025
I feel it's mostly about branding the research topic and having super neat figures (that being said I'm probably jealous knowing I will probably never have such a Nature paper 😬)
100
Paul Lerner @lernerp.bsky.social · 22/01/2025
I have the same impression for every other NLP/ML papers published in Nature (I read only a few)
110
Paul Lerner @lernerp.bsky.social · 22/01/2025
Really appreciate the feedback on this paper! It was mainly inspired by Valentin Hofmann et al. DagoBERT/"Superbizarre" papers
111
Paul Lerner @lernerp.bsky.social · 22/01/2025
Hope you enjoyed the presentation!
022
Reposted by Paul Lerner
Yoav Goldberg @yoavgo.bsky.social · 30/12/2024
RL promises "systems that can adapt to their environment". However, no RL system that I know of actually fulfill anything close to this goal, and, furthermore, I'd argue that all the current RL methodologies are actively hostile to this goal. Prove me wrong.
10534
Paul Lerner @lernerp.bsky.social · 08/01/2025
best prompt ever
010