Sign in

Maarten Sap

@maartensap.bsky.social
1.8K followers 220 following 44 posts

Working on #NLProc for social good. Currently at LTI at CMU. 🏳‍🌈

PostsRepliesMedia
Reposted by Maarten Sap
Akhila Yerukola @akhilayerukola.bsky.social · 01/10/2026
Did you know gifting four koi 🐟 is offensive in Japan (4 = "shi" = death 💀), or that chopsticks 🥢upright in rice is inappropriate in China (like funeral incense 🕯️)? Top VLMs don't! Our #COLM2026 paper introduces 🌏NormViz to measure visual norm understanding in 16 countries. Best model: 25.3% 🚨 🧵👇
Overview of NORMVIZ-BENCH for visual norm understanding, consisting of 3,272 contrastive image pairs across 16 countries that differ only in the culturally relevant behavior. (a) A contrastive pair example from Japan; group accuracy requires a model to correctly classify both images in a pair. (b) Images are sourced via Text-to-Image generation,
inpainting, and Retrieval. (c) Benchmark statistics and pair type distribution.
182
Reposted by Maarten Sap
Carolyn Rose, Kavčić-Moura Professor of Lang. Tech. and HCI, CMU @carolynrose.bsky.social · 19/09/2026
I am grateful to the ACL executive committee for nominating me for the position of ACL VP Elect.The ballot is now open for ACL members to vote. We are up against unprecedented challenges in the world and in our community, but if we join forces, we can move forward in synergy, strength, and success.
021
Maarten Sap @maartensap.bsky.social · 02/09/2026
Is the safety-capability tradeoff for LLMs real? Or could it be an artefact of the benchmarks we use to measure safety?? We did some explorations with psychometrics-inspired multi-dimensional IRT models and created BenchMIRT to explore these questions! See 🧵
0113
Maarten Sap @maartensap.bsky.social · 15/08/2026
I can't believe the first CMU Sapling Xuhui Zhou is now a Doctor! My first solo advisee is now a PhD!! HUGE Congrats @nlpxuhui.bsky.social , I'm so proud, and grateful to have been able to work together!
3240
Reposted by Maarten Sap
Giuseppe Attanasio @gattanasio.cc · 22/06/2026
Speech translation scores well on benchmarks, but is it actually usable when it counts? 🎙️ New paper: Ouvia, a user-centered framework for evaluating speech translation in real-world 1-to-1 communication. 1,700+ interactions, 174 speakers, 3 dialect groups. github.com/g8a9/ouvia #nlp #nlproc
1113
Reposted by Maarten Sap
Mingqian Zheng @mingqian-zheng.bsky.social · 13/05/2026
LLMs refuse ambiguous queries that look harmful but aren't. Can they recover once users clarify, while staying safe? Our new interactive multi-turn benchmark measures both. 🚨 Turns out: not both at once.
1113
Reposted by Maarten Sap
pluralistic-ai.bsky.social @pluralistic-ai.bsky.social · 30/03/2026
🚨We are excited to announce the 2nd Pluralistic Alignment Workshop at #ICML2026 in Seoul! Submissions: openreview.net/group?id=ICM... 🗓️Deadline: May 3 More details pluralistic-alignment.github.io We invite work on pluralistic alignment across technical, philosophical, and societal perspectives!
183
Reposted by Maarten Sap
Dr. Jordan Taylor @jordant.bsky.social · 10/03/2026
🎨💻 What is a “high-quality” or “aesthetic” image according to generative AI developers? Happy to share that our investigation of the LAION-Aesthetics Predictor has been accepted at #FAccT2026! 🧵 (1/5) Take a look at a preprint here: arxiv.org/abs/2601.09896
Screenshot of an academic paper titled "The Algorithmic Gaze of Image Quality Assessment: An Audit and Trace Ethnography of the LAION-Aesthetics Predictor" authored by Jordan Taylor, William Agnew, Maarten Sap, Sarah E. Fox, and Haiyi Zhu
1285
Maarten Sap @maartensap.bsky.social · 24/02/2026
Presenting two papers! On Tuesday at 12:40pm (Room I; 12pm session): 1-2-3 Check: Enhancing Contextual Privacy in LLM via Multi-Agent Reasoning arxiv.org/abs/2508.07667 (followed by poster session at 1pm)
arxiv.org
1-2-3 Check: Enhancing Contextual Privacy in LLM via Multi-Agent Reasoning
Addressing contextual privacy concerns remains challenging in interactive settings where large language models (LLMs) process information from multiple sources (e.g., summarizing meetings with private...
150
Reposted by Maarten Sap
Natalie Shapira @natalieshapira.bsky.social · 23/02/2026
You can read more in the full paper: www.researchgate.net/publication/... There is also an interactive web that contains logs of the authentic interactions: agentsofchaos.baulab.info
researchgate.net
(PDF) Agents of Chaos
PDF | We report an exploratory red-teaming study of autonomous language-model-powered agents deployed in a live laboratory environment with persistent... | Find, read and cite all the research you nee...
042
Reposted by Maarten Sap
Natalie Shapira @natalieshapira.bsky.social · 23/02/2026
In this amazing multidisciplinary collaboration, we report our early experience with the @openclaw-x.bsky.social ->
14022
Maarten Sap @maartensap.bsky.social · 22/02/2026
On my way to Paris for the IASEAI conference www.iaseai.org/our-programs...! Who will be there?
iaseai.org
IASEAI - International Association for Safe and Ethical AI
Building a global movement for safe and ethical AI. Join IASEAI to ensure AI systems operate safely and ethically, benefiting all of humanity.
010
Maarten Sap @maartensap.bsky.social · 02/02/2026
🚀 Apply to CMU LTI’s Summer 2026 “Language Technology for All” internship! 🎓 Open to pre‑doctoral students new to language tech (non‑CS backgrounds welcome). 🔬 12–14 weeks in‑person in Pittsburgh — travel + stipend paid. 💸 Deadline: Feb 20, 11:59pm ET. Apply → forms.gle/cUu8g6wb27Hs...
forms.gle
CMU LTI Summer 2026 Internship Program Application
We are looking for applicants for the Carnegie Mellon University Language Technology Institute's Summer 2026 "Language Technology for All" internship program. The main goal of this internship is to pr...
21512
Maarten Sap @maartensap.bsky.social · 13/01/2026
I'm excited to announce the Call for Papers for the Social Context (SoCon) and Integrating NLP and Psychology to Study Social Interactions (NLPSI) workshop, @ LREC '26 in Palma de Mallorca, Spain! 🗓Deadline: February 16, 2026 🌐Website: socon-nlpsi.github.io 🗓Workshop: May 12, 2026
socon-nlpsi.github.io
SoCon-NLPSI'26 | Home
Natural Language Processing (NLP) has undergone a significant evolution, opening up the possibility of capturing high-level aspects of human communication. Key areas of interest include the pragmatics...
0157
Maarten Sap @maartensap.bsky.social · 22/12/2025
I'm very excited about our new work which aims to model causes and effects on stories online! Narratives and stories are everywhere, so it's helpful to be able to understand how people use them in nuanced ways.
0143
Reposted by Maarten Sap
Mingqian Zheng @mingqian-zheng.bsky.social · 20/10/2025
How and when should LLM guardrails be deployed to balance safety and user experience? Our #EMNLP2025 paper reveals that crafting thoughtful refusals rather than detecting intent is the key to human-centered AI safety. 📄 arxiv.org/abs/2506.00195 🧵[1/9]
193
Reposted by Maarten Sap
Lorenzo Xiao @lrz-persona.bsky.social · 17/10/2025
📣📣 Announcing the first PersonaLLM Workshop on LLM Persona Modeling. If you work on persona driven LLMs, social cognition, HCI, psychology, cognitive science, cultural modeling, or evaluation, do not miss the chance to submit. Submit here: openreview.net/group?id=Neu...
openreview.net
NeurIPS 2025 Workshop Mexico City PersonaNLP
Welcome to the OpenReview homepage for NeurIPS 2025 Workshop Mexico City PersonaNLP
141
Maarten Sap @maartensap.bsky.social · 15/10/2025
I’m ✨ super excited and grateful ✨to announce that I'm part of the 2025 class of #PackardFellows (www.packard.org/2025fellows). The @packardfdn.bsky.social and this fellowship will allow me to explore exciting research directions towards culturally responsible and safe AI 🌍🌈
1111
Reposted by Maarten Sap
Kshitish Ghate @kghate.bsky.social · 14/10/2025
🚨New paper: Reward Models (RMs) are used to align LLMs, but can they be steered toward user-specific value/style preferences? With EVALUESTEER, we find even the best RMs we tested exhibit their own value/style biases, and are unable to align with a user >25% of the time. 🧵
1127
Maarten Sap @maartensap.bsky.social · 14/10/2025
Saplings take #COLM2025! Featuring Group lunch, amazing posters, and a panel with Yoshua Bengio!
1161
Reposted by Maarten Sap
Queer in AI @queerinai.com · 09/10/2025
We are launching our Graduate School Application Financial Aid Program (www.queerinai.com/grad-app-aid) for 2025-2026. We’ll give up to $750 per person to LGBTQIA+ STEM scholars applying to graduate programs. Apply at openreview.net/group?id=Que.... 1/5
queerinai.com
Grad App Aid — Queer in AI
179
Maarten Sap @maartensap.bsky.social · 06/10/2025
I'm also giving a talk at #COLM2025 Social Simulation workshop (sites.google.com/view/social-...) on Unlocking Social Intelligence in AI, at 2:30pm Oct 10th!
sites.google.com
Social Simulation with LLMs
Important Information Workshop Date: October 10, 2025 Workshop Room: 523AB Workshop Time Zone: Montreal (UTC -4) Contact Email: social-simulation@googlegroups.com
060
Maarten Sap @maartensap.bsky.social · 06/10/2025
Headed to #COLM2025 today! Here's five of our papers that were accepted, and when & where to catch them 👇
160
Reposted by Maarten Sap
Valentin Hofmann @valentinhofmann.bsky.social · 16/09/2025
📢 New #COLM2025 paper 📢 Standard benchmarks give every LLM the same questions. This is like testing 5th graders and college seniors with *one* exam! 🥴 Meet Fluid Benchmarking, a capability-adaptive eval method delivering lower variance, higher validity, and reduced cost. 🧵
34110
Maarten Sap @maartensap.bsky.social · 26/08/2025
That's a lot of people! Fall Sapling lab outing, welcoming our new postdoc Vasudha, and visitors Tze Hong and Chani! (just missing Jocelyn)
0120
Maarten Sap @maartensap.bsky.social · 25/08/2025
I'm excited cause I'm teaching/coordinating a new unique class, where we teach new PhD students all the "soft" skills of research, incl. ideation, reviewing, presenting, interviewing, advising, etc. Each lecture is taught by a different LTI prof! It takes a village! maartensap.com/11705/Fall20...
maartensap.com
LTI 11-705 Introduction to Research in Language Technologies (Fall 2025)
2312
Maarten Sap @maartensap.bsky.social · 22/08/2025
I spoke to Forbes about why model "welfare" is a silly framing to an important issue; models don't have feelings, and it's a big distraction from real questions like tensions between safety vs. user utility, which are NLP/HCI/policy questions www.forbes.com/sites/victor...
1143
Maarten Sap @maartensap.bsky.social · 20/08/2025
Super super excited about this :D :D
7270
Reposted by Maarten Sap
Language Technologies Institute | CMU @ltiatcmu.bsky.social · 27/06/2025
Hand gestures are a major mode of human communication, but they don't always translate well across cultures. New research from @akhilayerukola.bsky.social, @maartensap.bsky.social and others is aimed at giving AI systems a hand with overcoming cultural biases: lti.cmu.edu/news-and-eve...
lti.cmu.edu
Using Hand Gestures To Evaluate AI Biases - Language Technologies Institute - School of Computer Science - Carnegie Mellon University
LTI researchers have created a model to help generative AI systems understand the cultural nuance of gestures.
083
Reposted by Maarten Sap
Language Technologies Institute | CMU @ltiatcmu.bsky.social · 26/06/2025
New research from LTI, UMich, & Allen Institute for AI: LLMs don’t just hallucinate – sometimes, they lie. When truthfulness clashes with utility (pleasing users, boosting brands), models often mislead. @nlpxuhui.bsky.social and @maartensap.bsky.social discuss the paper: lti.cmu.edu/news-and-eve...
lti.cmu.edu
Does Your Chatbot Swear to Tell the Truth? - Language Technologies Institute - School of Computer Science - Carnegie Mellon University
New research finds that LLM-based agents can't always be trusted to be truthful
032
Reposted by Maarten Sap
Daniel Chechelnitsky @dchechel.bsky.social · 17/06/2025
What if AI played the role of your sassy gay bestie 🏳️‍🌈 or AAVE-speaking friend 👋🏾? You: “Can you plan a trip?” 🤖 AI: “Yasss queen! let’s werk this babe✨💅” LLMs can talk like us, but it shapes how we trust, rely on & relate to them 🧵 📣 our #FAccT2025 paper: bit.ly/3HJ6rWI [1/9]
1136
Reposted by Maarten Sap
Julia Mendelsohn @jmendelsohn2.bsky.social · 21/05/2025
📣 Super excited to organize the first workshop on ✨NLP for Democracy✨ at COLM @colmweb.org!! Check out our website: sites.google.com/andrew.cmu.e... Call for submissions (extended abstracts) due June 19, 11:59pm AoE #COLM2025 #LLMs #NLP #NLProc #ComputationalSocialScience
sites.google.com
NLP 4 Democracy - COLM 2025
14718
Reposted by Maarten Sap
Language Technologies Institute | CMU @ltiatcmu.bsky.social · 12/05/2025
Notice our new look? We're thrilled to unveil our new logo – representing our vision, values, and the future ahead. Stay tuned for more!
031
Maarten Sap @maartensap.bsky.social · 29/04/2025
super excited about this 🥰🥰
1200
Reposted by Maarten Sap
Xuhui Zhou @nlpxuhui.bsky.social · 28/04/2025
When interacting with ChatGPT, have you wondered if they would ever "lie" to you? We found that under pressure, LLMs often choose deception. Our new #NAACL2025 paper, "AI-LIEDAR ," reveals models were truthful less than 50% of the time when faced with utility-truthfulness conflicts! 🤯 1/
1259
Reposted by Maarten Sap
Neel Bhandari @neelbhandari.bsky.social · 17/04/2025
1/🚨 𝗡𝗲𝘄 𝗽𝗮𝗽𝗲𝗿 𝗮𝗹𝗲𝗿𝘁 🚨 RAG systems excel on academic benchmarks - but are they robust to variations in linguistic style? We find RAG systems are brittle. Small shifts in phrasing trigger cascading errors, driven by the complexity of the RAG pipeline 🧵
195
Maarten Sap @maartensap.bsky.social · 06/03/2025
RLHF is built upon some quite oversimplistic assumptions, i.e., that preferences between pairs of text are purely about quality. But this is an inherently subjective task (not unlike toxicity annotation) -- so we wanted to know, do biases similar to toxicity annotation emerge in reward models?
1243
Maarten Sap @maartensap.bsky.social · 26/02/2025
My PhD student Akhila's been doing some incredible cultural work in the last few years! Check out out latest work on cultural safety and hand gestures, showing most vision and/or language AI systems are very cross-culturally unsafe!
0273
Maarten Sap @maartensap.bsky.social · 21/02/2025
Super excited to unveil this work! LLMs need to ask better questions, and our method with synthetic data corruption can help generalize to other interesting LLM improvements (more to come on that ;) )
080
Reposted by Maarten Sap
Xuhui Zhou @nlpxuhui.bsky.social · 19/02/2025
LLM agents can code—but can they ask clarifying questions? 🤖💬 Tired of coding agents wasting time and API credits, only to output broken code? What if they asked first instead of guessing? 🚀 (New work led by Sanidhya Vijay: www.linkedin.com/in/sanidhya-...)
173
Maarten Sap @maartensap.bsky.social · 01/02/2025
Saplings Raclette night was a success 🥰
0200
Maarten Sap @maartensap.bsky.social · 07/01/2025
CMU LTI is hosting predoc interns this summer, centered around "Language Technologies for All"! Please apply and circulate! lti.cs.cmu.edu/news-and-eve...
lti.cs.cmu.edu
CMU LTI Language Technology for All Internship 2025 - Language Technologies Institute - School of Computer Science - Carnegie Mellon University
The LTI is currently seeking applicants for the summer 2025 Language Technology for All Internship
1198
Maarten Sap @maartensap.bsky.social · 07/12/2024
Headed to #TayNeurIPS in Vancouver! Message me on Whova if you're going and wanna meet up! #Neurips #Neurips2024
0100
Reposted by Maarten Sap
Language Technologies Institute | CMU @ltiatcmu.bsky.social · 20/11/2024
Looking for all your LTI friends on Bluesky? The LTI Starter Pack is here to help! go.bsky.app/NhTwCVb
6159
Maarten Sap @maartensap.bsky.social · 19/11/2024
Maria is an incredible researcher and mentor! I highly encourage applying!!
060
Reposted by Maarten Sap
Natalie Shapira @natalieshapira.bsky.social · 17/11/2024
A starter pack for researchers that are interested in #ToM #theory-of-mind and #Pragmatics #NLP #NLProc go.bsky.app/P6gL8vL
2114
Maarten Sap @maartensap.bsky.social · 16/11/2024
I heard this opinion a bunch through the conference but #EMNLP2024 was really good vibes! Congrats to the organizers 😁👏
3493
Maarten Sap @maartensap.bsky.social · 08/11/2024
Also, I'll be giving a talk on "Computational Methods of Social Causes and Effects of Stories" at the workshop on Narrative understanding on Nov 15th #emnlp2024 Excited to share some recent work with my students Joel & Jocelyn and collaborators @mariaa.bsky.social @andrewpiper.bsky.social!
1183
Maarten Sap @maartensap.bsky.social · 08/11/2024
Excited to attend EMNLP 2024 next week! A couple of my students are presenting their work!
181