Sign in

Anna Wegmann

@annawegmann.bsky.social
925 followers 428 following 52 posts

Postdoctoral Researcher at Utrecht University | Including different styles in NLP | she/her annawegmann.github.io

PostsRepliesMedia
Reposted by Anna Wegmann
Marco @mcognetta.bsky.social · 30/09/2026
🚨 [Token][ization] Paper Alert 🚨 Tokenization is a wildly understudied area of language modeling despite it having effects across all of NLP. Over the past ~8 months, 32 (!) tokenizer researchers put together the most comprehensive survey of the field. Check it out!
112333
Reposted by Anna Wegmann
INLG 2026 @inlg.bsky.social · 04/09/2026
🚂 The early registration for #INLG2026 departs September 15th. Standard fares after that. Utrecht, Oct 17–21. 2026.inlgmeeting.org/registration...
2026.inlgmeeting.org
INLG2026
The 19th International Natural Language Generation Conference is scheduled to be held in Utrecht, the Netherlands from October 17 to 21, 2026.
022
Reposted by Anna Wegmann
CS-NLP Group, Utrecht University @cs-nlp-uu.bsky.social · 04/09/2026
Check out this batch of papers headed straight to Budapest! #EMNLP2026 #NLProc #NLP #ComputationalLinguistics

Main Conference

A Survey on Representing Linguistic Style: Challenges and Opportunities
Anna Wegmann, Cristina Aggazzotti, Rafael Alberto Rivera Soto, Dong Nguyen

STEB: Style Text Embedding Benchmark (Main)
Rafael Alberto Rivera Soto, Anna Wegmann, Cristina Aggazzotti

Findings

A Circuit for Plural Reference: How LLMs Represent and Retrieve Singular and Plural Entities
Anh Dang, Rick Nouwen, Massimo Poesio

From Vision to Language: Investigating Causal Information Flow in Multimodal Decision-Making
Davide Testa, Hugh Mee Wong, Alessandro Lenci, Bernardo Magnini, Albert Gatt

Demo & Workshops

Neurobiber: Fast and Interpretable Linguistic Feature Extraction (Demo)
Kenan Alkiek, Anna Wegmann, Jian Zhu, David Jurgens

Acquire, Repair, Preserve: A Diagnosis-Guided Post-Training Recipe for Small-Model Dialogue Game Agents (LM Playschool)
Nan Li
131
Reposted by Anna Wegmann
INLG 2026 @inlg.bsky.social · 04/09/2026
🚂 127 #INLG2026 papers are on track for Utrecht! A sample: "Justify your prompts!", "Self-Verification is All You Need To Pass The Japanese Bar Examination", and live commentary generation for Age of Empires II. Congrats to all authors. Utrecht, Oct 17–21. 2026.inlgmeeting.org/accepted-pap...
2026.inlgmeeting.org
INLG2026
The 19th International Natural Language Generation Conference is scheduled to be held in Utrecht, the Netherlands from October 17 to 21, 2026.
042
Anna Wegmann @annawegmann.bsky.social · 01/09/2026
Nathan Lambert arguing for RLHF against the haters saying it's just style transfer -- youtu.be/o6l6tJQgUg4 on slide 50 around 35 minute mark
000
Anna Wegmann @annawegmann.bsky.social · 25/08/2026
3/3 papers accepted at #EMNLP2026! Thanks to great co-authors. More soon...
060
Anna Wegmann @annawegmann.bsky.social · 21/08/2026
Style is important quote n of ongoing [...] the most immediate problem to be solved in the attack on sociolinguistic structure is the quantification of the dimension of style. — William Labov (1972)
020
Reposted by Anna Wegmann
Julia Mendelsohn @jmendelsohn2.bsky.social · 28/07/2026
While I'm posting for #IC2S2, I want to share my new lab and website with you all! Our lab is called "Computational Analysis of Text and Society" (CATS) and I finally updated the website! Check it out: cats-group.github.io
cats-group.github.io
CATS Lab
CATS (Computational Analysis of Text and Society) is an interdisciplinary research lab at the University of Maryland.
1222
Reposted by Anna Wegmann
INLG 2026 @inlg.bsky.social · 16/07/2026
🆕 Check out our lineup of workshops! Both research and position papers are welcome. 2026.inlgmeeting.org/workshops-tu...
0108
Reposted by Anna Wegmann
Andrea Nini @andreanini.com · 13/07/2026
Our latest research (led by Sadie Barlow, with Dr Edoardo Manino) found LambdaG doesn’t need a calibration dataset for well-calibrated likelihood ratios, just a simple correction! This simplifies analysis and makes it the only authorship verification method without calibration. Paper on arXiv now.
andreanini.com
LambdaG: A simplified approach to Authorship Verification
In our latest research, led by Sadie Barlow and with Dr Edoardo Manino, we discovered a very important property of LambdaG: it does not require a calibration dataset to produce well calibrated likelihood ratios. A simple correction is sufficient. This means that the analysis is significantly simpler to implement and explain. LambdaG is therefore also the only existing authorship verification method that does not require a calibration analysis to produce likelihood ratios. The paper, currently under review, is available on arXiv here:
074
Reposted by Anna Wegmann
SIGGEN & INLG @siggen.bsky.social · 14/07/2026
❗️Direct submission deadline for #INLG2026 has been extended to July 18th (AoE). Be sure to proof-read your work, and good luck! #CfP #NLProc
053
Reposted by Anna Wegmann
INLG 2026 @inlg.bsky.social · 25/06/2026
The 2nd Call for Papers for #INLG2026 is out! Two ways to submit: 📄 Direct submissions, archival and not, by 15 July: openreview.net/group?id=acl... 💍 ARR commitments by 5 August: openreview.net/group?id=acl... ——— 📍Utrecht, NL — Oct 17–21, 2026, right before EMNLP 2026.inlgmeeting.org/calls.html
openreview.net
INLG 2026 Conference Direct Submissions
Welcome to the OpenReview homepage for INLG 2026 Conference Direct Submissions
087
Reposted by Anna Wegmann
CS-NLP Group, Utrecht University @cs-nlp-uu.bsky.social · 18/06/2026
Our heartfelt congratulations to Eduardo Calò, who successfully defended his PhD thesis titled "Automatically Expressing the Meaning of Logical Formulae in Natural Language" on 17 June at the Academiegebouw.
151
Reposted by Anna Wegmann
INLG 2026 @inlg.bsky.social · 13/05/2026
The Call for Papers for #INLG2026 is out! 🗓️ Submit by July 15 (AoE) 💍 ARR commit by August 5 🆕 Squibs welcomed (raising an issue without needing to solve it) 🆕 Non-archival track for work in progress ——— 📍Utrecht, NL — Oct 17–21, 2026, right before EMNLP 2026.inlgmeeting.org/calls.html
062
Reposted by Anna Wegmann
Paul Röttger @paul-rottger.bsky.social · 27/04/2026
New paper w/ UK AISI: Millions of people now use AI to help them write and communicate. In three experiments (14k participants, 3m+ human ratings) we show that AI writing assistance systematically distorts writer personas – their perceived beliefs, personality, and identity. 🧵
24413
Anna Wegmann @annawegmann.bsky.social · 22/04/2026
A dream come true. A conference in a train museum. What will the breaks be like? Choo Choo! 🚂🚂
020
Reposted by Anna Wegmann
Jenna Russell @jennarussell.bsky.social · 07/04/2026
Would you realize if the book you were reading was AI? What if it was humanized to remove AI-speak? We find that even without using stylistic cues (e.g., word choice or sentence structure) narrative choices alone give AI fiction away!
927675
Reposted by Anna Wegmann
Hye Sun Yun @hyesunyun.bsky.social · 08/04/2026
Patients ask LLMs medical questions — but how they phrase it matters more than it should. Our new preprint explores how different phrasings of patient health questions can lead to inconsistent conclusions, even with the same evidence. [1/6] Full Paper: arxiv.org/abs/2604.05051
2255
Reposted by Anna Wegmann
Anna Rogers @annarogers.bsky.social · 13/04/2026
📢 The workshop on Insights from negative results will be back at EMNLP'26! Your most-insightful failures can be submitted in 4 pages by June 25. It's also possible to commit short papers reviewed through ARR. insights-workshop.github.io/2026/cfp
insights-workshop.github.io
2026 Call for Papers
Workshop on Insights from Negative Results in NLP
03211
Anna Wegmann @annawegmann.bsky.social · 03/04/2026
Yay. Finally a conference next to my house! Hope to see many of the #NLProc #NLP community there 🤗
061
Reposted by Anna Wegmann
Marlene Lutz @marlutz.bsky.social · 12/03/2026
Survey-style tests developed for humans may not predict how LLMs actually behave. Our #EACL2026 paper shows they can even be misleading when measuring racism and sexism! Check out the paper 👇🏼
021
Reposted by Anna Wegmann
Mor Naaman @informor.bsky.social · 11/03/2026
🥁🥁🥁 Newly out from us today in Science Advances: “Biased AI Writing Assistants Shift Users’ Attitudes on Societal Issues”. Large Language Models are providing users with autocomplete writing suggestions on many platforms. Could these suggestions shift users’ own attitudes? (spoiler: YES) (1/7)
5195105
Reposted by Anna Wegmann
Daniel Paleka @dpaleka.bsky.social · 20/02/2026
Can LLMs figure out who you are from your anonymous posts? From a handful of comments, LLMs can infer where you live, what you do, and your interests; then search for you on the web. New 📄 w/ @SimonLermenAI, @joshua_swans, @AerniMichael, Nicholas Carlini, @florian_tramer 🧵
812442
Reposted by Anna Wegmann
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 13/02/2026
our paper on data mixing for LMs is out! while building Olmo 3, we saw gaps between data mixing literature and real practice 🐠choosing proxy size, # runs, sampling, regression, constraints.. 🐟data shifts during LM dev: can we reuse past experiments? Olmix tackles them all!
1294
Reposted by Anna Wegmann
bhyravajjula.bsky.social @bhyravajjula.bsky.social · 31/10/2025
If you're attending #EMNLP2025, we'll be presenting virtually in Gather Session 1 on Nov 5 at 4pm PT. Come say hello! w/ the wonderful: @mellymeldubs.bsky.social Anna Preus, @mariaa.bsky.social Paper: arxiv.org/abs/2510.16713 Code/Data: github.com/darthbhyrava/wisp Dash: poetry.darthbhyrava.com
181
Reposted by Anna Wegmann
David Jurgens @davidjurgens.bsky.social · 06/11/2025
What if a single model could recognize an author's writing style no matter what language they wrote in? 🌍✍️ Our new #EMNLP2025 paper explores multilingual authorship representation, showing how training across 36 languages can sharpen stylistic signals and reduce topic bias. 👇🧵
1182
Reposted by Anna Wegmann
Dong Nguyen @dongng.bsky.social · 04/11/2025
New opinion paper out with Esther Ploeger (Aalborg University): We Need to Measure Data Diversity in NLP — Better and Broader at #EMNLP2025 (main) aclanthology.org/2025.emnlp-m...
aclanthology.org
We Need to Measure Data Diversity in NLP — Better and Broader
Dong Nguyen, Esther Ploeger. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
2154
Anna Wegmann @annawegmann.bsky.social · 04/11/2025
Lot's of exciting work on linguistic style this year at #EMNLP2025 #EMNLP! Including work on machine-text detection, authorship representation and more 🧵 with anthology links below 📣 with an open call to everyone to add style work that's missing
191
Reposted by Anna Wegmann
Catherine Arnett @catherinearnett.bsky.social · 25/09/2025
I have a new blog post about the so-called “tokenizer-free” approach to language modeling and why it’s not tokenizer-free at all. I also talk about why people hate tokenizers so much!
45915
Anna Wegmann @annawegmann.bsky.social · 22/10/2025
I successfully defended my PhD in Dutch fashion and required a PhD certificate in Latin. Thank you to the amazing people that got me here, a.o. @dongng.bsky.social and the ones I blur here.
1342
Reposted by Anna Wegmann
Indira Sen @indiiigo.bsky.social · 16/10/2025
Come join next Wednesday if you want to rant about society's love-hate relationship with LLMs!
0137
Anna Wegmann @annawegmann.bsky.social · 20/08/2025
Is this the Dutch budget cuts or does utrecht uni really not want me to come to the office? My highlight is the door that has been broken for weeks, with the only change being a laminated piece of paper saying I should enter uni maze through two other buildings.
100
Reposted by Anna Wegmann
Ruben van de Vijver @rubenvandevijver.bsky.social · 01/08/2025
Tussen Mönchengladbach en Venlo rijden geen treinen. De dienstregeling wordt gehandhaafd door een bus. De bus codeswitcht
Tussen Mönchengladbach en Venlo rijden geen treinen. De dienstregeling wordt gehandhaafd door een bus. De bus codeswitcht: een monitor waarop staat „de bus hält“.
0183
Anna Wegmann @annawegmann.bsky.social · 05/08/2025
Utrecht is back from #ACL2025! We had a blast. I should have posted this before but here are some papers from people in our group that were presented at ACL.
130
Reposted by Anna Wegmann
Craig Schmidt @craigschmidt.com · 30/07/2025
I'm sadly not at #ACL2025, but the work on tokenization seem to continue to explode. Here are the tokenization related papers I could find, in no particular order. Let me know if I missed any.
2124
Anna Wegmann @annawegmann.bsky.social · 29/07/2025
Since people at #ACL2025 are very interested in tokenization, a reminder to join the discussion on discord set up by @mcognetta.bsky.social
092
Anna Wegmann @annawegmann.bsky.social · 28/07/2025
Anyone tried the kiss the cook lunch place at #ACL2025?
000
Anna Wegmann @annawegmann.bsky.social · 28/07/2025
I will present our #ACL2025 paper Tokenization is Sensitive to Language Variation in the poster session after Tuesday's keynote, 10.30 - 12.00 in Hall 4/5
0100
Reposted by Anna Wegmann
Tiago Pimentel @tpimentel.bsky.social · 27/07/2025
@philipwitti.bsky.social will be presenting our paper "Tokenisation is NP-Complete" at #ACL2025 😁 Come to the language modelling 2 session (Wednesday morning, 9h~10h30) to learn more about how challenging tokenisation can be!
062
Reposted by Anna Wegmann
Tiago Pimentel @tpimentel.bsky.social · 27/07/2025
We are presenting this paper at #ACL2025 😁 Find us at poster session 4 (Wednesday morning, 11h~12h30) to learn more about tokenisation bias!
0112
Anna Wegmann @annawegmann.bsky.social · 27/07/2025
Im at #ACL2025 this week. Happy to chat about measuring linguistic style, data diversity, creating synthetic data for analyzing (L)LMs, authorship attribution, paraphrases and tokenizers. Let’s chat if you’re around
010
Reposted by Anna Wegmann
Maria Antoniak @mariaa.bsky.social · 17/07/2025
The #ACL2025 #ACL2025NLP feed is up and running! It matches both hashtags and any posts from or mentions of @aclmeeting.bsky.social Pin it to your home 📌 and enjoy! bsky.app/profile/did:...
24814
Reposted by Anna Wegmann
Matthias Orlikowski @morlikow.bsky.social · 24/07/2025
Who's presenting on subjectivity in annotation (human label variation, learning from disagreement, perspectivism) at #ACL2025? papers by e.g. @liweijiang.bsky.social @tiancheng.bsky.social @gabriellalapesa.bsky.social @romanklinger.de keynote @verenarieser.bsky.social link to full list below ⤵️
2122
Anna Wegmann @annawegmann.bsky.social · 24/07/2025
I love it.
020
Reposted by Anna Wegmann
Ece Takmaz @ecekt.bsky.social · 24/07/2025
I'll be attending ACL 2025 in Vienna! Looking forward to seeing people there!😊🇦🇹 We are going to present 'LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks' aclanthology.org/2025.acl-sho... #acl2025 #acl2025nlp
0102
Anna Wegmann @annawegmann.bsky.social · 22/07/2025
How come the @aclmeeting.bsky.social underline page was set to release July 20 last Friday and now promises access only on the 24th? Access to papers and videos remains evasive less than a week before the conference.
Screenshot of a text: This event page is still under construction and will be released on July 24, 2025.
100
Reposted by Anna Wegmann
Indira Sen @indiiigo.bsky.social · 21/07/2025
Do LLMs represent the people they're supposed simulate or provide personalized assistance to? We review the current literature in our #ACL2025 Findings paper and investigating what researchers conclude about the demographic representativeness of LLMs: osf.io/preprints/so... 1/
Screenshot of our paper "Missing the Margins: A Systematic Literature Review on the Demographic Representativeness of LLMs"Details about what we annotated in our systematic review
2249
Reposted by Anna Wegmann
Julia Mendelsohn @jmendelsohn2.bsky.social · 22/07/2025
#ic2s2 I’ll be talking about this paper in one hour in Vingen 1+2!
081
Reposted by Anna Wegmann
Anders Giovanni Møller @handle.invalid · 22/07/2025
#ic2s2 I’ll have a poster (#21) today on 𝐭𝐡𝐞 𝐢𝐦𝐩𝐚𝐜𝐭 𝐨𝐟 𝐠𝐞𝐧𝐞𝐫𝐚𝐭𝐢𝐯𝐞 𝐀𝐈 𝐨𝐧 𝐬𝐨𝐜𝐢𝐚𝐥 𝐦𝐞𝐝𝐢𝐚 ✌🏼
0123
Reposted by Anna Wegmann
Miriam Schirmer @miriamschirmer.bsky.social · 22/07/2025
Join my talk on #ChildObjectification on TikTok at #IC2S2 today at 2:30 PM (📍 Social Good & Ethics). I built classification models to detect objectifying language and found: 10% of comments refer to appearance, 3% are objectifying. Models struggle with this task, with #RoBERTa outperforming GPT-4.
Topic Overview of Comments on Videos with Children
3102