Sign in

David Mortensen

@davidrmortensen.bsky.social
804 followers 1.2K following 75 posts

I make colorless green GPUs sleep brrriously. Computational phonology, morphology, language change models, speech/language technologies (especially for people with disabilities).

PostsRepliesMedia
Reposted by David Mortensen
Amanda Bertsch @abertsch.bsky.social · 07/11/2025
Can LLMs accurately aggregate information over long, information-dense texts? Not yet… We introduce Oolong, a dataset of simple-to-verify information aggregation questions over long inputs. No model achieves >50% accuracy at 128K on Oolong!
Performance of a sweep of models on Oolong-synth and Oolong-real. Performance decreases with increasing context length, sometimes steeply.
35020
Reposted by David Mortensen
Kshitish Ghate @kghate.bsky.social · 14/10/2025
🚨New paper: Reward Models (RMs) are used to align LLMs, but can they be steered toward user-specific value/style preferences? With EVALUESTEER, we find even the best RMs we tested exhibit their own value/style biases, and are unable to align with a user >25% of the time. 🧵
1127
Reposted by David Mortensen
Andy Liu @andyliu.bsky.social · 02/10/2025
🚨New Paper: LLM developers aim to align models with values like helpfulness or harmlessness. But when these conflict, which values do models choose to support? We introduce ConflictScope, a fully-automated evaluation pipeline that reveals how models rank values under conflict. (📷 xkcd)
1164
Reposted by David Mortensen
Niyati Bafna @niyatibafna.bsky.social · 04/07/2025
🔈When LLMs solve tasks with a mid-to-low resource input or target language, their output quality is poor. We know that. But can we put our finger on what breaks inside the LLM? We introduce the 💥 translation barrier hypothesis 💥 for failed multilingual generation with LLMs. arxiv.org/abs/2506.22724
2267
Reposted by David Mortensen
Valentin Hofmann @valentinhofmann.bsky.social · 09/05/2025
Thrilled to share that this is out in @pnas.org today! 🎉 We show that linguistic generalization in language models can be due to underlying analogical mechanisms. Shoutout to my amazing co-authors @weissweiler.bsky.social, @davidrmortensen.bsky.social, Hinrich Schütze, and Janet Pierrehumbert!
1356
Reposted by David Mortensen
Lindia Tjuatja @lindiatjuatja.bsky.social · 09/06/2025
When it comes to text prediction, where does one LM outperform another? If you've ever worked on LM evals, you know this question is a lot more complex than it seems. In our new #acl2025 paper, we developed a method to find fine-grained differences between LMs: 🧵1/9
27020
Reposted by David Mortensen
Syeda Nahida Akter @reasyaay.bsky.social · 01/05/2025
RL boosts LLM reasoning—but why stop at math & code? 🤔 Meet Nemotron-CrossThink—a method to scale RL-based self-learning across law, physics, social science & more. 🔥Resulting in a model that reasons broadly, adapts dynamically, & uses 28% fewer tokens for correct answers! 🧵↓
153
Reposted by David Mortensen
Verena Blaschke @verenablaschke.bsky.social · 29/04/2025
On my way to #NAACL2025 where I'll give a keynote at the noisy text workshop (WNUT), presenting some of the challenges & methods for dialect NLP + also discussing dialect speakers' perspectives! 🗨️ Beyond “noisy” text: How (and why) to process dialect data 🗓️ Saturday, May 3, 9:30–10:30
1267
Reposted by David Mortensen
Kshitish Ghate @kghate.bsky.social · 29/04/2025
Excited to announce our #NAACL2025 Oral paper! 🎉✨ We carried out the largest systematic study so far to map the links between upstream choices, intrinsic bias, and downstream zero-shot performance across 131 CLIP Vision-language encoders, 26 datasets, and 55 architectures!
1216
Reposted by David Mortensen
Kwanghee Choi @juice500ml.bsky.social · 29/04/2025
Can self-supervised models 🤖 understand allophony 🗣? Excited to share my new #NAACL2025 paper: Leveraging Allophony in Self-Supervised Speech Models for Atypical Pronunciation Assessment arxiv.org/abs/2502.07029 (1/n)
21510
Reposted by David Mortensen
Nishant Subramani @ ACL @nsubramani23.bsky.social · 29/04/2025
🚀 Excited to share a new interp+agents paper: 🐭🐱 MICE for CATs: Model-Internal Confidence Estimation for Calibrating Agents with Tools appearing at #NAACL2025 This was work done @msftresearch.bsky.social last summer with Jason Eisner, Justin Svegliato, Ben Van Durme, Yu Su, and Sam Thomson 1/🧵
1128
Reposted by David Mortensen
Xuhui Zhou @nlpxuhui.bsky.social · 28/04/2025
When interacting with ChatGPT, have you wondered if they would ever "lie" to you? We found that under pressure, LLMs often choose deception. Our new #NAACL2025 paper, "AI-LIEDAR ," reveals models were truthful less than 50% of the time when faced with utility-truthfulness conflicts! 🤯 1/
1259
Reposted by David Mortensen
Neel Bhandari @neelbhandari.bsky.social · 17/04/2025
1/🚨 𝗡𝗲𝘄 𝗽𝗮𝗽𝗲𝗿 𝗮𝗹𝗲𝗿𝘁 🚨 RAG systems excel on academic benchmarks - but are they robust to variations in linguistic style? We find RAG systems are brittle. Small shifts in phrasing trigger cascading errors, driven by the complexity of the RAG pipeline 🧵
195
Reposted by David Mortensen
Chise @sailorrooscout.bsky.social · 31/03/2025
THIS IS HUGE! Researchers at McMaster University have discovered a NEW peptide antibiotic that targets a broad range of disease-causing bacteria INCLUDING those RESISTANT to existing antibiotics. This discovery marks the first potential new class of antibiotics in NEARLY 30 YEARS. 🧪🧵⬇️
22792272731
Reposted by David Mortensen
Naomi Saphra @nsaphra.bsky.social · 27/03/2025
Life update: I'm starting as faculty at Boston University @bucds.bsky.social in 2026! BU has SCHEMES for LM interpretability & analysis, I couldn't be more pumped to join a burgeoning supergroup w/ @najoung.bsky.social @amuuueller.bsky.social. Looking for my first students, so apply and reach out!
CDS building which looks like a jenga tower
3524213
David Mortensen @davidrmortensen.bsky.social · 19/03/2025
You should read Article 1 of the United States Constitution. It's a trip.
010
David Mortensen @davidrmortensen.bsky.social · 19/03/2025
There can be only one DB joke. And that is DB.
010
Reposted by David Mortensen
Johann-Mattis List @lingulist.de · 17/03/2025
New preprint by @annikatjuka.bsky.social, Robert Forkel, Christoph Rzymski, and myself available, presenting a new version of the Database of Cross-Linguistic Colexifications (CLICS). "Advancing the Database of Cross-Linguistic Colexifications with New Workflows and Data" arxiv.org/abs/2503.11377
arxiv.org
Advancing the Database of Cross-Linguistic Colexifications with New Workflows and Data
Lexical resources are crucial for cross-linguistic analysis and can provide new insights into computational models for natural language learning. Here, we present an advanced database for comparative ...
173
Reposted by David Mortensen
Connor Ewing @cmewing.bsky.social · 16/03/2025
Finally found a way to shorten faculty meetings.
1725659
Reposted by David Mortensen
Nick Fleisher @nickfleisher.bsky.social · 12/03/2025
No student anywhere in America has said something as antisemitic as this
112421
Reposted by David Mortensen
dchiang.bsky.social @dchiang.bsky.social · 08/03/2025
The meeting will feature keynote addresses by @mohitbansal.bsky.social, @davidrmortensen.bsky.social, Karen Livescu, and Heng Ji. Plus all of your great talks and posters! nlp.nd.edu/msld25
nlp.nd.edu
Midwest Speech and Language Days 2025
041
Reposted by David Mortensen
Erin Jean Warde @erinjeanwarde.bsky.social · 06/03/2025
I’ve been thinking about this reading from Isaiah 58 since I heard it at the Ash Wednesday service today. “Is not this the fast that I choose: to loose the bonds of injustice, to undo the thongs of the yoke, to let the oppressed go free, and to break every yoke?
818833
Reposted by David Mortensen
Chris Hayes @chrislhayes.bsky.social · 06/03/2025
“Again, the mice used for clinical purposes did not undergo gender transition.” www.rollingstone.com/politics/pol...
rollingstone.com
Trump Decried Millions Spent 'Making Mice Transgender.' It Was Cancer and Asthma Research
President Trump falsely claimed that Biden spent $8 million on 'making mice transgender,' but the real research was for human health.
51959431163
Reposted by David Mortensen
U.S. Representative Al Green @algreen.house.gov · 06/03/2025
Today, the House GOP censured me for speaking out for the American people against @POTUS’s plan to cut Medicaid. I accept the consequences of my actions, but I refuse to stay silent in the face of injustice. #WeShallOvercome x.com/repalgreen/s...
x.com
Congressman Al Green on X: "Today, the House GOP censured me for speaking out for the American people against @POTUS’s plan to cut Medicaid. I accept the consequences of my actions, but I refuse to stay silent in the face of injustice. #WeShallOvercome https://t.co/sVklRmPCJl" / X
Today, the House GOP censured me for speaking out for the American people against @POTUS’s plan to cut Medicaid. I accept the consequences of my actions, but I refuse to stay silent in the face of injustice. #WeShallOvercome https://t.co/sVklRmPCJl
1024010592919207
Reposted by David Mortensen
Joel Mire @joelmire.bsky.social · 06/03/2025
Reward models for LMs are meant to align outputs with human preferences—but do they accidentally encode dialect biases? 🤔 Excited to share our paper on biases against African American Language in reward models, accepted to #NAACL2025 Findings! 🎉 Paper: arxiv.org/abs/2502.12858 (1/10)
Screenshot of Arxiv paper title, "Rejected Dialects: Biases Against African American Language in Reward Models," and author list: Joel Mire, Zubin Trivadi Aysola, Daniel Chechelnitsky, Nicholas Deas, Chrysoula Zerva, and Maarten Sap.
13811
David Mortensen @davidrmortensen.bsky.social · 05/03/2025
I read a paper about search, but I can't quite remember what it's called.
191
Reposted by David Mortensen
Danny To Eun Kim @teknology.bsky.social · 05/03/2025
🚨New Breakthrough in Tip-of-the-Tongue (TOT) Retrieval Research! We address data limitations and offer a fresh evaluation method for these complex queries. Curious how TREC TOT track test queries are created? Check out this thread 🧵 and our paper 📄: arxiv.org/abs/2502.17776
arxiv.org
Tip of the Tongue Query Elicitation for Simulated Evaluation
Tip-of-the-tongue (TOT) search occurs when a user struggles to recall a specific identifier, such as a document title. While common, existing search systems often fail to effectively support TOT scena...
2188
Reposted by David Mortensen
Robert Evans (the Only Robert Evans) @iwriteok.bsky.social · 04/03/2025
everything is so shitty, read this story about a genuinely good man who saw he had an opportunity to save millions of lives and threw himself into doing so. the world is full of heroes like him.
6183521802
Reposted by David Mortensen
This skeet will self destruct @pardoguerra.bsky.social · 04/03/2025
I humbly put this forward as a possible campaign for the new electric VW
10638120
Reposted by David Mortensen
CALL TO ACTIVISM @calltoactivism.bsky.social · 02/03/2025
This is Marco Rubio explaining how the USA promised to defend Ukraine forever if they got rid of their nuclear arsenal left after the Soviet Union fell. This is why lil marco was sinking into the couch. He was hoping we wouldn’t find it…so don’t RT right now this very second.
20255395436868
Reposted by David Mortensen
Alt18F @alt18f.bsky.social · 01/03/2025
18F was doing exactly the type of work that DOGE claims to want – yet we were eliminated shortly after midnight. Read our letter to the American people: 18f.org
18f.org
We're not done yet | 18F
697187396805
David Mortensen @davidrmortensen.bsky.social · 28/02/2025
There is a truism that all languages are equally complex, but they differ in how they are complex. There is some evidence for this (we have a paper about phonotactics and word length in Dutch varieties that is consistent with this). However, this is still an open question.
020
David Mortensen @davidrmortensen.bsky.social · 28/02/2025
Certainly, the writing systems for Hindi and Gujarati are straightforward. However, the interaction between case, gender, and animacy in Hindi is *not* simple. Plenty of languages grant no grammatical status to any of these three things. They're not necessary for communication, but Hindi likes them.
110
David Mortensen @davidrmortensen.bsky.social · 28/02/2025
Very good, but unlike Wittgenstein, I don't hit students.
100
Reposted by David Mortensen
Mark Riedl @markriedl.bsky.social · 27/02/2025
What could possibly go wro… oh why do I bother?
2133
Reposted by David Mortensen
Naomi Saphra @nsaphra.bsky.social · 28/02/2025
wow cool. your findings really remind me of my paper that has a citation count juuust below my h-index
0262
David Mortensen @davidrmortensen.bsky.social · 28/02/2025
When languages are viewed as social games (rather than coding systems for content), their labyrinthine corridors can be seen as essential game elements. When we try to assign all these evolutionary spandrels functional motivations, we force ourselves to be intellectually dishonest. 3/3
140
David Mortensen @davidrmortensen.bsky.social · 28/02/2025
Instead of demonstrating your superiority through simple and elegant means (e.g., pistols at dawn) you choose to do it by manipulating three kinds of markers, five decks of cards, two forms of currency, and turns with three or five stages depending on the current stage in the game. 2/3
120
David Mortensen @davidrmortensen.bsky.social · 28/02/2025
A common observation among students in my Subword Modeling course is that language is unnecessarily complicated. At some level, this is true. Languages are like those complicated table-top games I don't like. 1/3
370
Reposted by David Mortensen
Niyati Bafna @niyatibafna.bsky.social · 27/02/2025
Dialects lie on continua of (structured) linguistic variation, right? And we can’t collect data for every point on the continuum...🤔 📢 Check out DialUp, a technique to make your MT model robust to the dialect continua of its training languages, including unseen dialects. arxiv.org/abs/2501.16581
1135
Reposted by David Mortensen
Pranjal @pranjal2041.bsky.social · 26/02/2025
What if AI agents did software engineering like humans—seeing the screen & using any developer tool? Introducing Programming with Pixels: an SWE environment where agents control VSCode via screen perception, typing & clicking to tackle diverse tasks. programmingwithpixels.com 🧵
184
Reposted by David Mortensen
Akhila Yerukola @akhilayerukola.bsky.social · 26/02/2025
Did you know? Gestures used to express universal concepts—like wishing for luck—vary DRAMATICALLY across cultures? 🤞means luck in US but deeply offensive in Vietnam 🚨 📣 We introduce MC-SIGNS, a test bed to evaluate how LLMs/VLMs/T2I handle such nonverbal behavior! 📜: arxiv.org/abs/2502.17710
Figure showing that interpretations of gestures vary dramatically across regions and cultures. ‘Crossing your fingers,’ commonly used in the US to wish for good luck, can be deeply offensive to female audiences in parts of Vietnam. Similarly, the 'fig gesture,' a playful 'got your nose' game with children in the US, carries strong sexual connotations in Japan and can be highly offensive.
1337
Reposted by David Mortensen
Daniel Benneworth-Gray @danielgray.com · 24/02/2025
oh cool I’ve been meaning to learn anxiety
221126411651
Reposted by David Mortensen
Tris @trisresists.bsky.social · 25/02/2025
Just going to leave this here should it become useful to anyone…
5794139419610
Reposted by David Mortensen
Mark Riedl @markriedl.bsky.social · 22/02/2025
There are a bunch of published papers referencing “vegetative electron microscopy” because of Generative AI retractionwatch.com/2025/02/10/v...
2378