Sign in

Paul Lerner

@lernerp.bsky.social
50 followers 56 following 50 posts

Research Scientist @ NuMind paullerner.github.io

PostsRepliesMedia
Paul Lerner @lernerp.bsky.social · 28/08/2026
I'm starting this new personal project: autollm, a python library for easy data distillation of LLMs for text classification, information extraction, and open-ended tasks github.com/PaulLerner/a...
github.com
120
Paul Lerner @lernerp.bsky.social · 04/06/2026
Hi folks, I'm searching for a Research Scientist position in Paris starting from September! Let me know if you hear about an opportunity :)
000
Paul Lerner @lernerp.bsky.social · 14/05/2026
Great discussion at LREC :)
000
Reposted by Paul Lerner
Inria Paris NLP (ALMAnaCH team) @inriaparisnlp.bsky.social · 08/05/2026
Parallel Corpora of Scholarly Documents for English-French Machine Translation Ziqian Peng, Lichao Zhu, Rachel Bawden, Maud Bénard, Éric de la Clergerie, Mathilde Huguin, Natalie Kübler, Paul Lerner, Alexandra Mestivier & François Yvon 📅 11th May | 16:30–16:54 | BUCC (remote)
101
Reposted by Paul Lerner
MLIA ISIR @mlia-isir.bsky.social · 06/05/2026
➡️ Can Multimodal LLMs Generate Pedagogical Questions? Thomas Gerald, Sahar Ghannay, Julie Lascar, Paul Lerner @lernerp.bsky.social, Anne Vilnat in collaboration with LISN @lisnlab.bsky.social x.com/LISNLAB 📅 Thurs., 14 May, 11:00 - 12:40 (long, poster)
021
Reposted by Paul Lerner
MLIA ISIR @mlia-isir.bsky.social · 06/05/2026
➡️ Assessing the Political Fairness of Multilingual LLMs: A Case Study based on a 21-way Multiparallel EuroParl Dataset Paul Lerner @lernerp.bsky.social, François Yvon @yvofr.bsky.social 📅 Wed., 13 May, 11:40 (long, oral) 📖 arxiv.org/abs/2510.20508
021
Reposted by Paul Lerner
MLIA ISIR @mlia-isir.bsky.social · 06/05/2026
➡️ Parallel Corpora of Scholarly Documents for English-French Machine Translation #BUCC Z. Peng, @lichaozhu.bsky.social @rachelbawden.bsky.social @maudbenard.bsky.social Éric de la Clergerie, @mathildehuguin.bsky.social @nataliekubler.bsky.social @lernerp.bsky.social A. Mestivier & @yvofr.bsky.social
122
Paul Lerner @lernerp.bsky.social · 05/05/2026
Happy to chat if you're at LREC next week, I'll be presenting Wednesday at 11:40 in Session O4 "Evaluation, Validation, Quality Assurance and Benchmarking Methodologies"
000
Reposted by Paul Lerner
Inria Paris NLP (ALMAnaCH team) @inriaparisnlp.bsky.social · 03/04/2026
We are very happy to announce our next seminar: Paul Lerner @lernerp.bsky.social (ISIR, Sorbonne Université & CNRS) "Controlling Linguistic Variability in Large Language Models" on Friday 10th April 2026, 11am CET. Details here 👉 almanach.inria.fr/seminars-en....
ALMAnaCH seminar: Paul Lerner, “Controlling Linguistic Variability in Large Language Models”, 10/04/2026
132
Reposted by Paul Lerner
Leonie Weissweiler @weissweiler.bsky.social · 11/12/2025
🧑‍🔬I’m recruiting PhD students in Natural Language Processing @unileipzig.bsky.social Computer Science, together with @scadsai.bsky.social! Topics include, but aren’t limited to: 🔎Linguistic Interpretability 🌍Multilingual Evaluation 📖Computational Typology Please share! #NLProc #NLP
14225
Reposted by Paul Lerner
MLIA ISIR @mlia-isir.bsky.social · 24/11/2025
The team meeting of the week was presented by Alexandre Vérine, from PSL, about "Quality and Diversity in generative models through the lens of f-divergences." Thanks a lot for this interesting talk!
001
Reposted by Paul Lerner
MLIA ISIR @mlia-isir.bsky.social · 28/10/2025
Accepted to a Workshop (1/2): "Self-Retrieval from Distant Contexts for Document-Level Machine Translation", accepted to the Conference on Machine Translation (WMT25), from @ziqianpeng.bsky.social, @rachelbawden.bsky.social, @yvofr.bsky.social
102
Paul Lerner @lernerp.bsky.social · 06/11/2025
Come work with @yvofr.bsky.social @weissweiler.bsky.social and me at @mlia-isir.bsky.social for a M2 internship on Assessing the Morphological Competence of LLMs! For 5-6 months from February or March 2026. Paid 600€/month
132
Paul Lerner @lernerp.bsky.social · 24/10/2025
What's the plural of "LLM-as-a-Judge"?
000
Paul Lerner @lernerp.bsky.social · 23/10/2025
We find that LLMs translate some political parties unfairly using a new version of EuroParl, fully multi-parallel and including (political) metadata hal.science/hal-05328251
hal.science
Assessing the Political Fairness of Multilingual LLMs: A Case Study based on a 21-way Multiparallel EuroParl Dataset
The political biases of Large Language Models (LLMs) are usually assessed by simulating their answers to English surveys. In this work, we propose an alternative framing of political biases, relying on principles of fairness in multilingual translation. We systematically compare the translation quality of speeches in the European Parliament (EP), observing systematic differences with majority parties from left, center, and right being better translated than outsider parties. This study is made possible by a new, 21-way multiparallel version of EuroParl, the parliamentary proceedings of the EP, which includes the political affiliations of each speaker. The dataset consists of 1.5M sentences for a total of 40M words and 249M characters. It covers three years, 1000+ speakers, 7 countries, 12 EU parties, 25 EU committees, and hundreds of national parties.
110
Paul Lerner @lernerp.bsky.social · 15/10/2025
introducing 🤔 ppllm, a Python Library to Compute LLM's Perplexity and Surprisal github.com/PaulLerner/p...
github.com
GitHub - PaulLerner/ppllm: 🤔 A Python Library to Compute LLM's Perplexity and Surprisal
🤔 A Python Library to Compute LLM's Perplexity and Surprisal - PaulLerner/ppllm
100
Paul Lerner @lernerp.bsky.social · 04/09/2025
make.org/FR/consultat...
make.org
Contribute to current consultations - Comment l’IA peut-elle améliorer la vie des Français en limitant les risques ? - Make.org
Finding proposals is easier when working together. Discover a democratic place where you can discuss the big issues you care about, submit your proposals concerning them and vote on proposals proposed...
000
Reposted by Paul Lerner
Pasquale Minervini @neuralnoise.com · 05/07/2025
"in 2025 we will have flying cars" 😂😂😂
839991
Paul Lerner @lernerp.bsky.social · 07/07/2025
Last week, I presented my work on "Assessing the Political Biases of Multilingual LLMs" at the EALM workshop @ TALN 2025 ! Thanks again to the ANR Diké project for organizing the workshop
100
Reposted by Paul Lerner
MLIA ISIR @mlia-isir.bsky.social · 10/06/2025
📢 🎉 The team has one paper accepted to #MTsummit2025! "Investigating Length Issues in Document-level Machine Translation" by @ziqianpeng.bsky.social, @rachelbawden.bsky.social and @yvofr.bsky.social in collaboration with @inriaparisnlp.bsky.social 📍 Geneva | 🗓️ 23-27,June 📕 arxiv.org/abs/2412.17592
arxiv.org
Investigating Length Issues in Document-level Machine Translation
Transformer architectures are increasingly effective at processing and generating very long chunks of texts, opening new perspectives for document-level machine translation (MT). In this work, we chal...
111
Paul Lerner @lernerp.bsky.social · 16/06/2025
"meticulously" is so absent from this list (from aclanthology.org/2025.coling-... )
000
Paul Lerner @lernerp.bsky.social · 16/06/2025
Am I the only reviewer that actually fills this "Reviewer Checklist"? And why do Area Chairs never answer when the paper needs to be desk-rejected? And reviews are due in 3 days 🫠
000
Reposted by Paul Lerner
MLIA ISIR @mlia-isir.bsky.social · 10/06/2025
For the EALM Workshop "On Assessing the Political Biases of Multilingual Large Language Models" by @lernerp.bsky.social Laurène Cave, @haldaume3.bsky.social Léo Labat, Gaël Lejeune, Pierre-Antoine Lequeu, @bpiwowar.bsky.social Nazanin Shafiabadi and yvofr.bsky.social, collaborated with the STIH lab
102
Paul Lerner @lernerp.bsky.social · 20/02/2025
Amazed at what a COLING paper could look like in the 80's
010
Paul Lerner @lernerp.bsky.social · 11/02/2025
Hope you enjoyed our poster at #AISummit! I'm standing next to Pierre-Antoine Lequeu, @salimhafid.bsky.social, and @manonberriche.bsky.social but there's more people involved! Zoom-in to read their names or learn more about the project here about.make.org/democratic-c...
031
Paul Lerner @lernerp.bsky.social · 25/01/2025
accurate quote for NLP researchers visiting Louvre Abu Dhabi after COLING 2025
010
Paul Lerner @lernerp.bsky.social · 22/01/2025
Really appreciate the feedback on this paper! It was mainly inspired by Valentin Hofmann et al. DagoBERT/"Superbizarre" papers
111
Paul Lerner @lernerp.bsky.social · 22/01/2025
Hope you enjoyed the presentation!
022
Reposted by Paul Lerner
Yoav Goldberg @yoavgo.bsky.social · 30/12/2024
RL promises "systems that can adapt to their environment". However, no RL system that I know of actually fulfill anything close to this goal, and, furthermore, I'd argue that all the current RL methodologies are actively hostile to this goal. Prove me wrong.
10534
Paul Lerner @lernerp.bsky.social · 08/01/2025
best prompt ever
010
Paul Lerner @lernerp.bsky.social · 06/01/2025
We offer a M2 internship on Visual Question Answering at LISN (Paris-Saclay University, co-supervised by Thomas Gerald, Sahar Ghannay, and Anne Vilnat) The goal is to create a dataset of questions for education (based on schoolbook content) Duration of 5 or 6 months (starting in March or April)
111
Paul Lerner @lernerp.bsky.social · 23/12/2024
if you're interested in morphological segmentation, I started working on a word-based/lexematic approach where a given word is segmented into a base and affix (e.g. "invaluable" -> "in- valuable" where "valuable" can be further decomposed into "value -able") github.com/PaulLerner/n...
github.com
GitHub - PaulLerner/neoseg: A tool for Lexematic Segmentation by Paul Lerner
A tool for Lexematic Segmentation by Paul Lerner. Contribute to PaulLerner/neoseg development by creating an account on GitHub.
020
Reposted by Paul Lerner
Yoav Goldberg @yoavgo.bsky.social · 17/12/2024
but, their summaries are still far better than pretty much all we had before, only that now we are stuck because there is no way of improving the quality in a principled (or even non-principled) way. so kinda stuck, not a great position to be in.
1354
Paul Lerner @lernerp.bsky.social · 16/12/2024
My brother did an awesome PhD 🤩
000
Paul Lerner @lernerp.bsky.social · 16/12/2024
I got two papers accepted at #COLING2025 🤓
120
Paul Lerner @lernerp.bsky.social · 27/11/2024
@jkamps.bsky.social Hi, I'm looking for this great keynote you gave at TALN-CORIA 2023 in Paris, entitled something like "Has NLP become IR? Or has IR become NLP?" Any link or reference on this topic?
000
Paul Lerner @lernerp.bsky.social · 20/11/2024
Hey ML/NLP teachers out there: during practical works, how do you deal with the fact that training/fine-tuning a model takes time? Like do your students sometime have to wait 20 minutes or more while the model is training? If not, how do you design your practical work so they're training-free?
000
Paul Lerner @lernerp.bsky.social · 15/11/2024
You might have followed me on twitter previously @_lerner_paul
010