Sign in

Elinor

@elinorpd.bsky.social
1.6K followers 453 following 248 posts

PhD @ MIT CSAIL // researching LLM societal impacts & alignment previously @ MIT media lab, mila quebec / mcgill i like language and dogs and plants and ultimate frisbee and baking and sunsets. she/her elinorp-d.github.io

PostsRepliesMedia
Reposted by Elinor
Leshem (Legend) Choshen @EMNLP @lchoshen.bsky.social · 04/09/2026
She's a sexy fox 🦊 היא חתולה סקסית 😻 People believe translation is solved, because we dropped the bar. But the above is something no translation system (aka LLM) is close to doing. We made worldwide challenges and keep collecting so you can read use it join the authors...
181
Reposted by Elinor
Isabelle Augenstein @iaugenstein.bsky.social · 03/09/2026
Our paper on measuring epistemic diversity in LLMs is accepted to #EMNLP2026! We find that diversity has improved, though LLM output is much less diverse than Web search, especially for non-English. Paper: arxiv.org/pdf/2510.04226 Code/data: github.com/dwright37/ll... #NLProc @copenlu.bsky.social
1263
Reposted by Elinor
Shomir Wilson @shomir.bsky.social · 24/08/2026
I'm pleased to share a preprint of our accepted EMNLP main conference paper "Affective Context Amplifies Sycophancy in LLM Responses". Sycophancy, an unnatural level of agreement or praise for a person's ideas, is a problem for chatbots, sometimes with disastrous outcomes. arxiv.org/abs/2608.212...
arxiv.org
Affective Context Amplifies Sycophancy in LLM Responses
As conversational companions, large language models (LLMs) often have access to users' emotional states. We study how this affective context modulates LLM sycophancy in subjective, evaluative interact...
0141
Reposted by Elinor
Lucy Li @lucy3.bsky.social · 07/08/2026
Spent a good amount of this summer digging through a 20+ person team of teachers' free-text annotations, learning what "scaffolding" and "push for rigor" means, and iterating on this pipeline. We find that AI tutors, by default, frequently over-scaffold and rarely push for rigor.
27412
Reposted by Elinor
Henry Yuen @henryyuen.bsky.social · 01/08/2026
I spent many hours, days, nights, weekends in cafes, in my office, at home, trying to understand Ran Raz's classical parallel repetition theorem and whether I could prove a quantum version of it. This period of struggle was important for me.
1835
Elinor @elinorpd.bsky.social · 23/07/2026
Pluralistic alignment is thriving as a research agenda yet failing at its goal: making the AI systems people actually use more pluralistic🌈 🚨New position paper: we argue adoption in deployed models should be the fields main goal & we provide a roadmap of how to get there 🧵1/
3212
Reposted by Elinor
David Jurgens @davidjurgens.bsky.social · 13/07/2026
With my ACL PC abilities, I wanted to characterize who these folks are, whether they're actually "outsiders", whether their papers are much less likely to be accepted, is this growth just LLM slop? The whole analsis is here medium.com/@jurgens_245...
medium.com
Is the ACL Rolling Review actually broken?
ARR is broken! ARR is being overwhelmed with papers from outsiders! LLM-generated slop is ruining our peer review! There are not enough…
12811
Reposted by Elinor
Jonathan Stray @jonathanstray.com · 23/10/2025
GreenEarth is creating open source AI-driven recommender infrastructure for BlueSky. Type a prompt, see your feed change. We are here for the users, the builders, the dreamers. Join us. greenearthsocial.substack.com/p/introducin...
greenearthsocial.substack.com
Introducing GreenEarth
We're building advanced open source algorithms for social media
1010427
Reposted by Elinor
iNaturalist @inaturalist.bsky.social · 06/07/2026
We're looking for a CV/ML Engineer to help us improve the machine learning systems that power iNaturalist's species identification and geographic range modeling. If you're excited to help build tools that help millions of people engage with nature, we'd love to hear from you! Apply: buff.ly/YZqaW6c
Image of flowers with text overlaid saying: "iNaturalist: We're hiring! Computer vision/machine learning engineer. Full time, remote in the United States."
03921
Elinor @elinorpd.bsky.social · 14/06/2026
Just uploaded a new plant blog post on my website 🥰🪴🌸
270
Reposted by Elinor
Ehud Reiter @ehudreiter.bsky.social · 08/06/2026
New blog: I am worried by NLP research culture NLG and NLP are mostly much better in 2026 than when I got my PhD in 1990. Unfortunately research culture has gotten *worse” in this period, which really worries me as I retire. ehudreiter.com/2026/06/08/n...
ehudreiter.com
I am worried by NLP research culture
In most ways NLG and NLP are much better in 2026 than when I got my PhD in 1990. Unfortunately research culture has gotten *worse” in this period, which really worries me as I retire. We have…
1144
Elinor @elinorpd.bsky.social · 06/06/2026
excited to share that i'll be pursuing my phd in computer science at @csail.mit.edu starting this fall 🥳 🎓 i'm so grateful to be coadvised by the literal dream team: Jacob Andreas, @mbakker.bsky.social, and @mitchellgordon.bsky.social 🙌
4240
Reposted by Elinor
Benno Krojer @bennokrojer.bsky.social · 05/06/2026
My first last-author paper is out! If you saw this dog below and someone showed you the second image, would you consider them the same word/concept? (more examples in Ada's thread) We study if VLMs agree with humans on this and revisit old questions around shape vs. texture bias in vision
1124
Reposted by Elinor
Jessica Hullman @jessicahullman.bsky.social · 28/05/2026
Thoughts on metascientific consequences of AI-generated slides & ideas diluting the impression that speakers are commited to what they present. Science runs on personal attachment more than we admit. If it were a cake mix, how wouldn we add back an egg? statmodeling.stat.columbia.edu/2026/05/28/w...
statmodeling.stat.columbia.edu
What if scientists really were dispassionate observers, communicating ideas without irrational commitment? Look here, says AI. | Statistical Modeling, Causal Inference, and Social Science
1296
Reposted by Elinor
Ilias Chalkidis @kiddothe2b.bsky.social · 06/05/2026
📢 Our paper 🤯🧠 "Brainrot: Deskilling and Addiction are Overlooked AI Risks" has been accepted at the ACM Fairness, Accountability & Transparency (FAccT) conference 2026. The preprint is available: 👉 arxiv.org/abs/2605.03512 TL;DR 🧵 follows 👇 1/5
arxiv.org
Brainrot: Deskilling and Addiction are Overlooked AI Risks
The scope of AI safety and alignment work in generative artificial intelligence (GenAI) has so far mostly been limited to harms related to: (a) discrimination and hate speech, (b) harmful/inappropriat...
1215
Reposted by Elinor
Joachim Baumann @joachimbaumann.bsky.social · 01/05/2026
Can you boost your AI review scores by asking an LLM to rewrite your paper? Yes! We call it paper laundering Our @icmlconf.bsky.social spotlight paper argues current AI reviewers aren't ready to automate peer review, and outlines what a science of peer review automation should look like 🧵👇 #ICML2026
First page of the ICML 2026 spotlight paper "Stop Automating Peer Review Without Rigorous Evaluation" by Joachim Baumann, Jiaxin Pei, Sanmi Koyejo, and Dirk Hovy (Stanford University and Bocconi University). The abstract argues that today's AI systems should not be used to produce paper reviews, grounded in two empirical findings: a "hivemind effect" where AI reviewers show excessive agreement and reduce perspective diversity, and "paper laundering," where prompting an LLM to rewrite a paper trivially increases AI reviewer scores through stylistic changes rather than scientific improvements. The paper calls for a science of peer review automation rather than wholesale deployment of general-purpose LLMs.
44314
Reposted by Elinor
Petter Törnberg @pettertornberg.com · 01/05/2026
LLMs have been widely reported as left-wing biased. The finding has shaped policy and debate — with Trump banning "Woke AI". Our new paper challenges this story. It's not that the models are biased. It's that they think the auditor is. 🧵 w/ michelleschimmel.bsky.social‬arxiv.org/pdf/2604.27633
415750
Reposted by Elinor
Shahan Ali Memon @shahanmemon.bsky.social · 30/04/2026
Our position paper, "AI Welfare Is Bullshit" just got accepted to ICML 2026 @icmlconf.bsky.social! The AI welfare agenda has already begun to attract institutional investment at organizations such an @anthropic.com. We argue that this idea is essentially Frankfurtian Bullshit.
56219
Reposted by Elinor
Mark Riedl @markriedl.bsky.social · 30/04/2026
Designating that most uses of the term "goblin" and "gremlin" are not legitimate while designating most uses of the word "frog" as legitimate is the imposition of a set of values on a technology. Why do a small number of people in Silicon Valley get to decide whether goblins are inappropriate?
5736
Reposted by Elinor
Myra Cheng @myra.bsky.social · 23/04/2026
Patrick Sui and I are hosting an #ICLR2026 social for anyone with background/interest in the humanities! Room 210, 12-1:30pm on Friday 24 April!! Humanities-adjacent, humanities-curious, everyone is welcome! Should be a fun group! :)
1217
Elinor @elinorpd.bsky.social · 22/04/2026
Flying out of Boston to Brazil rn means I’m surrounded by poster tubes (ICLR attendees) and blue athletic gear w medals (Boston marathoners). It’s a cool crowd
050
Elinor @elinorpd.bsky.social · 21/04/2026
I'll be presenting OvertonBench at #ICLR2026 in Rio later this week! 📍Sat, Apr 25, 10:30am in Pavilion 4 (#4109) Please DM me if you'd like to chat about pluralistic / value alignment, societal impacts, epistemology, fairness, evals, etc
081
Reposted by Elinor
MIT Siegel Family Quest for Intelligence @mit-sqi.bsky.social · 17/04/2026
Congratulations to Jacob Andreas, was named a 2026 Edgerton Award recipient! The award recognizes exceptional teaching, research, and service at MIT! Prof. Andreas co-leads our Language and Thought Mission, and he is a dedicated and creative researcher and educator. news.mit.edu/2026/jacob-a...
news.mit.edu
Jacob Andreas and Brett McGuire named Edgerton Award winners
MIT associate professors Jacob Andreas and Brett McGuire have been selected as the winners of the 2026 Harold E. Edgerton Faculty Achievement Award for exceptional contributions to teaching, research,...
171
Elinor @elinorpd.bsky.social · 18/04/2026
My Hoya Rebecca put out the cutest teeny lil flowers today and I’m obsessed
030
Reposted by Elinor
David Bau @davidbau.bsky.social · 08/04/2026
Tech industry mottos have a mixed track record. But we should hold idealists to their ideals. And we should celebrate when they come through. The Mythos non-release is a remarkable moment of conviction. Thoughts: davidbau.com/archives/20... Bravo to Anthropic's "race the top".
1133
Reposted by Elinor
Gillian Hadfield @ghadfield.bsky.social · 03/04/2026
Democracy isn't a rulebook. It runs on daily interactions where people comply with norms and hold each other accountable. AI agents are about to join that system. We need to build them to read it. New paper with Rakshit Trivedi and Dylan Hadfield-Menell.
knightcolumbia.org
Building AI for the Democratic Matrix: A Technical Research Agenda for Normative Competence and Normative Institutions
To maintain democratic resilience, it is essential to build AI agents capable of choosing behaviors that mirror those of the human agents that constitute human democracies.
4112
Reposted by Elinor
Leshem (Legend) Choshen @EMNLP @lchoshen.bsky.social · 30/03/2026
Rebuttal season is here, yey🤞 With many asking me, I compiled the most common misconceptions Hope the tips help 🧵 All tips: docs.google.com/document/d/14Wax8M5w8F_8miDlYJ9-I6wqpelxlXjCEUbkNzNMqqE/edit?tab=t.0#heading=h.rfq27f356vmm #AI 🤖📈🧠
152
Reposted by Elinor
Mor Naaman @informor.bsky.social · 23/03/2026
New important (I hope) resource for academics working in this area.
191
Reposted by Elinor
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 15/03/2026
For a recent lab meeting, I wrote up a grab bag of ways to think about your development as a researcher during a PhD: emerge-lab.github.io/papers/an-un... Sharing in case folks find it useful or have feedback!
emerge-lab.github.io
610112
Reposted by Elinor
Mor Naaman @informor.bsky.social · 11/03/2026
🥁🥁🥁 Newly out from us today in Science Advances: “Biased AI Writing Assistants Shift Users’ Attitudes on Societal Issues”. Large Language Models are providing users with autocomplete writing suggestions on many platforms. Could these suggestions shift users’ own attitudes? (spoiler: YES) (1/7)
5195105
Elinor @elinorpd.bsky.social · 10/03/2026
There's been a lot of excitement about pluralistic value alignment 🌈 — AI that reflects the full range of human perspectives But no formal way to benchmark whether we're actually making progress. 🤔 Introducing 𝐎𝐕𝐄𝐑𝐓𝐎𝐍𝐁𝐄𝐍𝐂𝐇. 🎉Accepted to #ICLR2026 1/n 🧵
1231
Reposted by Elinor
Leshem (Legend) Choshen @EMNLP @lchoshen.bsky.social · 09/03/2026
Do LLMs Benefit from Their Own Words?🤔 In multi-turn chats, models are typically given their own past responses as context. But do their own words always help… Or are they more often a waste of compute and a distraction? 🧵 arxiv.org/abs/2602.24287
2374
Reposted by Elinor
Lucy Li @lucy3.bsky.social · 03/03/2026
Models are now expert math solvers, and so AI for math education is receiving increasing attention. Our new preprint evaluates 11 VLMs on our QA benchmark, DrawEduMath. We highlight a startling gap: models perform less well on inputs from K-12 students who need more help. 🧵
Title, author list, and two figures from the paper. 
Title: The Aftermath of DrawEduMath: Vision Language Models
Underperform with Struggling Students and Misdiagnose Errors
Authors: Li Lucy, Albert Zhang, Nathan Anderson, Ryan Knight, Kyle Lo
Figure 1: On the left is a math problem, where students are asked to draw x < 5/2 on a number line. The right side shows two example student responses that differ in correctness. DrawEduMath pairs each math problem with one student response, and prompts VLMs to answer questions about the student response.
Figure 2: VLMs consistently perform worse on answering DrawEduMath benchmark questions pertaining to erroneous student responses. Performance on non-erroneous student responses is labeled with specific VLMs’ names; that same model’s performance on erroneous student responses is directly below.
43512
Reposted by Elinor
Alexandra Olteanu @aolteanu.bsky.social · 03/03/2026
Yesterday was my last day at MSR. We recently learned that our roles were eliminated, and with them our little FATE Montreal team. I joined MSR a bit over 7.5 years ago while on active chemotherapy, and being at MSR has overlapped with so much change in my life.
4346
Reposted by Elinor
Emma Pierson @emmapierson.bsky.social · 26/02/2026
Our paper, "What's in My Human Feedback", received an oral presentation at ICLR! Our method automatically+interpretably identifies preferences in human feedback data; we use this to improve personalization + safety. Reach out if you have data/use cases to apply this to! arxiv.org/pdf/2510.26202
0283
Reposted by Elinor
Benno Krojer @bennokrojer.bsky.social · 11/02/2026
Finally we do test it empirically: finding some models where the embedding matrix of the LLM already provides decently interpretable nearest neighbors But this was not the full story yet... @mariusmosbach.bsky.social and @elinorpd.bsky.social nudged me to use contextual embeddings
111
Elinor @elinorpd.bsky.social · 11/02/2026
Really cool new work with surprising results! Highly recommend checking out the demo 👀
030
Reposted by Elinor
David Rand @dgrand.bsky.social · 04/02/2026
Grok fact-checks our paper on Grok fact-checking - and it approves!
1297
Reposted by Elinor
Shaily @shaily99.bsky.social · 02/02/2026
🎭 How do LLMs (mis)represent culture? 🧮 How often? 🧠 Misrepresentations = missing knowledge? spoiler: NO! At #CHI2026 we are bringing ✨TALES✨ a participatory evaluation of cultural (mis)reps & knowledge in multilingual LLM-stories for India 📜 arxiv.org/abs/2511.21322 1/10
14722
Elinor @elinorpd.bsky.social · 30/01/2026
this is amazing! made quick NYC & boston posters
030
Elinor @elinorpd.bsky.social · 30/01/2026
Potato is a great platform for researchers! Highly recommend (plus a great development team behind it)
010
Reposted by Elinor
Hanna Wallach @hannawallach.bsky.social · 29/01/2026
Microsoft Research NYC is hiring a researcher in the space of AI and society!
26240
Reposted by Elinor
David Bau @davidbau.bsky.social · 26/01/2026
What should academics be doing right now? I have been writing up some thoughts on what the research says about effective action, and what universities specifically can do. davidbau.github.io/poetsandnurs... It's on GitHub. Suggestions and pull requests welcome. github.com/davidbau/poe...
Federal agents with weapons drawn, moments before murdering American citizens on the streets of Minneapolis at the dawn of 2026.
03715
Elinor @elinorpd.bsky.social · 23/01/2026
I'll be presenting this work on January 25th (Hall 2, poster 41) at #AAAI2026 in Singapore! Please stop by and reach out if you'd like to chat 😁
030
Elinor @elinorpd.bsky.social · 23/01/2026
🎉 Excited to share our new paper which was accepted to #AAAI2026! As LLMs become increasingly used as sources of factual knowledge, we ask: Do they perform equitably across users of different backgrounds? 🧵⬇️ 1/6
121
Reposted by Elinor
Angelina Wang @angelinawang.bsky.social · 12/12/2025
Most LLM evals use API calls or offline inference, testing models in a memory-less silo. Our new Patterns paper shows this misses how LLMs actually behave in real user interfaces, where personalization and interaction history shape responses: arxiv.org/abs/2509.19364
13811
Reposted by Elinor
arXiv cs.AI Artificial Intelligence @csai-bot.bsky.social · 02/12/2025
Elinor Poole-Dayan, Jiayi Wu, Taylor Sorensen, Jiaxin Pei, Michiel A. Bakker: Benchmarking Overton Pluralism in LLMs arxiv.org/abs/2512.01351 arxiv.org/pdf/2512.01351 arxiv.org/html/2512.01351
021
Reposted by Elinor
Tomer Ullman @tomerullman.bsky.social · 20/11/2025
1212
Reposted by Elinor
Gautam Kamath @gautamkamath.com · 19/11/2025
Thoughtful (as always) blog post from Nicholas Carlini. "Are large language models worth it?" A nice read giving his perspective on risks of ML models. Post: nicholas.carlini.com/writing/2025... For people who prefer, this is the video of the talk from @colmweb.org www.youtube.com/watch?v=PngH...
13411
Reposted by Elinor
Avijit Ghosh @evijit.io · 13/11/2025
Extremely thrilled to talk about our new paper: "Who Evaluates AI’s Social Impacts? Mapping Coverage And Gaps In First And Third Party Evaluations". This is the first big project output from the @eval-eval.bsky.social coalition! Thread below:
1207