Sign in

Elinor

@elinorpd.bsky.social
1.6K followers 453 following 248 posts

PhD @ MIT CSAIL // researching LLM societal impacts & alignment previously @ MIT media lab, mila quebec / mcgill i like language and dogs and plants and ultimate frisbee and baking and sunsets. she/her elinorp-d.github.io

PostsRepliesMedia
Reposted by Elinor
Adam Visokay @avisokay.bsky.social · 02/10/2026
Looking forward to the Text as Data conference on Monday at UC Berkeley! I will be presenting our new paper "Multi-Perspective LLM Annotations for Valid Analyses in Subjective Tasks" which is a joint effort with Navya Mehrotra and @gligoric.bsky.social. 🧵
1224
Reposted by Elinor
Leshem (Legend) Choshen @EMNLP @lchoshen.bsky.social · 04/09/2026
She's a sexy fox 🦊 היא חתולה סקסית 😻 People believe translation is solved, because we dropped the bar. But the above is something no translation system (aka LLM) is close to doing. We made worldwide challenges and keep collecting so you can read use it join the authors...
181
Reposted by Elinor
Isabelle Augenstein @iaugenstein.bsky.social · 03/09/2026
Our paper on measuring epistemic diversity in LLMs is accepted to #EMNLP2026! We find that diversity has improved, though LLM output is much less diverse than Web search, especially for non-English. Paper: arxiv.org/pdf/2510.04226 Code/data: github.com/dwright37/ll... #NLProc @copenlu.bsky.social
1263
Reposted by Elinor
Shomir Wilson @shomir.bsky.social · 24/08/2026
I'm pleased to share a preprint of our accepted EMNLP main conference paper "Affective Context Amplifies Sycophancy in LLM Responses". Sycophancy, an unnatural level of agreement or praise for a person's ideas, is a problem for chatbots, sometimes with disastrous outcomes. arxiv.org/abs/2608.212...
arxiv.org
Affective Context Amplifies Sycophancy in LLM Responses
As conversational companions, large language models (LLMs) often have access to users' emotional states. We study how this affective context modulates LLM sycophancy in subjective, evaluative interact...
0141
Reposted by Elinor
Lucy Li @lucy3.bsky.social · 07/08/2026
Spent a good amount of this summer digging through a 20+ person team of teachers' free-text annotations, learning what "scaffolding" and "push for rigor" means, and iterating on this pipeline. We find that AI tutors, by default, frequently over-scaffold and rarely push for rigor.
27412
Reposted by Elinor
Henry Yuen @henryyuen.bsky.social · 01/08/2026
I spent many hours, days, nights, weekends in cafes, in my office, at home, trying to understand Ran Raz's classical parallel repetition theorem and whether I could prove a quantum version of it. This period of struggle was important for me.
1835
Elinor @elinorpd.bsky.social · 23/07/2026
Huge thanks to my amazing coauthors 🙏 Jillian Fisher @atoosakz.bsky.social Jacob Andreas @mitchellgordon.bsky.social @mbakker.bsky.social , and shoutout to @taylor-sorensen.bsky.social for all the helpful feedback! Paper: impactful-pluralistic-alignment.github.io 10/10
impactful-pluralistic-alignment.github.io
A Roadmap to Impactful Pluralistic Alignment Research
030
Elinor @elinorpd.bsky.social · 23/07/2026
Adoption will ultimately need frontier labs' cooperation. But much of the gap is ours (researchers) to close: making the empirical case, settling what ideal pluralistic behavior looks like, and building evals + methods feasible for deployment. The roadmap is in the paper 🗺️ 9/
110
Elinor @elinorpd.bsky.social · 23/07/2026
We consider the strongest objections: Shouldn't regulation handle this? What about open models🐳? Might pluralism already emerge on its own? Is this the wrong goal entirely🎯? We respond to 8 alternative views. Most, taken seriously, support our same call to action. 8/
120
Elinor @elinorpd.bsky.social · 23/07/2026
3: evals + methods aren't adoptable today. Existing metrics aren't hill-climbable, eg coverage scores reward adding perspectives even when it hurts accuracy or usefulness, and tradeoffs go unmeasured We need to build towards release-ready evals + methods fit for deployment. 7/
120
Elinor @elinorpd.bsky.social · 23/07/2026
2: we haven’t settled when pluralism is warranted or what it looks like ideally. "Subjective" queries are too vague. Research can inform how user+societal context *should* influence operationalization. We need a concrete goal that a model spec could state. 6/
110
Elinor @elinorpd.bsky.social · 23/07/2026
So why no adoption? Reason 1: the case for pluralism is mostly normative. No study has tested whether pluralistic responses actually benefit users: depolarization, epistemic autonomy, representation, etc. We call for human studies measuring benefits and costs together. 5
120
Elinor @elinorpd.bsky.social · 23/07/2026
Evaluations are where priorities show. No frontier lab publicly focuses on pluralism model card evaluations The one dedicated pluralism eval ever published (Muse Spark) measured individual preference prediction, and was dropped again for the 1.1 release. 4/
110
Elinor @elinorpd.bsky.social · 23/07/2026
We audit the public docs of the most widely used AI models: Anthropic, OpenAI, Google, xAI, Meta (+ top open-weight families) In behavior docs: 4 of 5 prescribe some multi-perspective behavior. None name pluralism as a goal or indicate training for it explicitly 3/
120
Elinor @elinorpd.bsky.social · 23/07/2026
To me, working on this shouldn't be just an academic pursuit. Pluralistic alignment aims to solve important problems in the real world, affecting billions of humans worldwide🌍 So our ask is collective: let's prioritize research that builds towards deployed impact 💪 2/
120
Elinor @elinorpd.bsky.social · 23/07/2026
Pluralistic alignment is thriving as a research agenda yet failing at its goal: making the AI systems people actually use more pluralistic🌈 🚨New position paper: we argue adoption in deployed models should be the fields main goal & we provide a roadmap of how to get there 🧵1/
3212
Reposted by Elinor
David Jurgens @davidjurgens.bsky.social · 13/07/2026
With my ACL PC abilities, I wanted to characterize who these folks are, whether they're actually "outsiders", whether their papers are much less likely to be accepted, is this growth just LLM slop? The whole analsis is here medium.com/@jurgens_245...
medium.com
Is the ACL Rolling Review actually broken?
ARR is broken! ARR is being overwhelmed with papers from outsiders! LLM-generated slop is ruining our peer review! There are not enough…
12811
Reposted by Elinor
Jonathan Stray @jonathanstray.com · 23/10/2025
GreenEarth is creating open source AI-driven recommender infrastructure for BlueSky. Type a prompt, see your feed change. We are here for the users, the builders, the dreamers. Join us. greenearthsocial.substack.com/p/introducin...
greenearthsocial.substack.com
Introducing GreenEarth
We're building advanced open source algorithms for social media
1010427
Reposted by Elinor
iNaturalist @inaturalist.bsky.social · 06/07/2026
We're looking for a CV/ML Engineer to help us improve the machine learning systems that power iNaturalist's species identification and geographic range modeling. If you're excited to help build tools that help millions of people engage with nature, we'd love to hear from you! Apply: buff.ly/YZqaW6c
Image of flowers with text overlaid saying: "iNaturalist: We're hiring! Computer vision/machine learning engineer. Full time, remote in the United States."
03921
Elinor @elinorpd.bsky.social · 14/06/2026
elinorp-d.github.io/blog/2026/ho...
elinorp-d.github.io
Hoya Rebecca | Elinor Poole-Dayan
The first hoya I bought when I moved to Cambridge to begin grad school at MIT
010
Elinor @elinorpd.bsky.social · 14/06/2026
Just uploaded a new plant blog post on my website 🥰🪴🌸
270
Reposted by Elinor
Ehud Reiter @ehudreiter.bsky.social · 08/06/2026
New blog: I am worried by NLP research culture NLG and NLP are mostly much better in 2026 than when I got my PhD in 1990. Unfortunately research culture has gotten *worse” in this period, which really worries me as I retire. ehudreiter.com/2026/06/08/n...
ehudreiter.com
I am worried by NLP research culture
In most ways NLG and NLP are much better in 2026 than when I got my PhD in 1990. Unfortunately research culture has gotten *worse” in this period, which really worries me as I retire. We have…
1144
Elinor @elinorpd.bsky.social · 06/06/2026
thank you! (and thanks to them!)
000
Elinor @elinorpd.bsky.social · 06/06/2026
thank you!!
000
Elinor @elinorpd.bsky.social · 06/06/2026
thank you!
000
Elinor @elinorpd.bsky.social · 06/06/2026
excited to share that i'll be pursuing my phd in computer science at @csail.mit.edu starting this fall 🥳 🎓 i'm so grateful to be coadvised by the literal dream team: Jacob Andreas, @mbakker.bsky.social, and @mitchellgordon.bsky.social 🙌
4240
Reposted by Elinor
Benno Krojer @bennokrojer.bsky.social · 05/06/2026
My first last-author paper is out! If you saw this dog below and someone showed you the second image, would you consider them the same word/concept? (more examples in Ada's thread) We study if VLMs agree with humans on this and revisit old questions around shape vs. texture bias in vision
1124
Reposted by Elinor
Jessica Hullman @jessicahullman.bsky.social · 28/05/2026
Thoughts on metascientific consequences of AI-generated slides & ideas diluting the impression that speakers are commited to what they present. Science runs on personal attachment more than we admit. If it were a cake mix, how wouldn we add back an egg? statmodeling.stat.columbia.edu/2026/05/28/w...
statmodeling.stat.columbia.edu
What if scientists really were dispassionate observers, communicating ideas without irrational commitment? Look here, says AI. | Statistical Modeling, Causal Inference, and Social Science
1296
Reposted by Elinor
Ilias Chalkidis @kiddothe2b.bsky.social · 06/05/2026
📢 Our paper 🤯🧠 "Brainrot: Deskilling and Addiction are Overlooked AI Risks" has been accepted at the ACM Fairness, Accountability & Transparency (FAccT) conference 2026. The preprint is available: 👉 arxiv.org/abs/2605.03512 TL;DR 🧵 follows 👇 1/5
arxiv.org
Brainrot: Deskilling and Addiction are Overlooked AI Risks
The scope of AI safety and alignment work in generative artificial intelligence (GenAI) has so far mostly been limited to harms related to: (a) discrimination and hate speech, (b) harmful/inappropriat...
1215
Reposted by Elinor
Joachim Baumann @joachimbaumann.bsky.social · 01/05/2026
Can you boost your AI review scores by asking an LLM to rewrite your paper? Yes! We call it paper laundering Our @icmlconf.bsky.social spotlight paper argues current AI reviewers aren't ready to automate peer review, and outlines what a science of peer review automation should look like 🧵👇 #ICML2026
First page of the ICML 2026 spotlight paper "Stop Automating Peer Review Without Rigorous Evaluation" by Joachim Baumann, Jiaxin Pei, Sanmi Koyejo, and Dirk Hovy (Stanford University and Bocconi University). The abstract argues that today's AI systems should not be used to produce paper reviews, grounded in two empirical findings: a "hivemind effect" where AI reviewers show excessive agreement and reduce perspective diversity, and "paper laundering," where prompting an LLM to rewrite a paper trivially increases AI reviewer scores through stylistic changes rather than scientific improvements. The paper calls for a science of peer review automation rather than wholesale deployment of general-purpose LLMs.
44314
Reposted by Elinor
Petter Törnberg @pettertornberg.com · 01/05/2026
LLMs have been widely reported as left-wing biased. The finding has shaped policy and debate — with Trump banning "Woke AI". Our new paper challenges this story. It's not that the models are biased. It's that they think the auditor is. 🧵 w/ michelleschimmel.bsky.social‬arxiv.org/pdf/2604.27633
415750
Reposted by Elinor
Shahan Ali Memon @shahanmemon.bsky.social · 30/04/2026
Our position paper, "AI Welfare Is Bullshit" just got accepted to ICML 2026 @icmlconf.bsky.social! The AI welfare agenda has already begun to attract institutional investment at organizations such an @anthropic.com. We argue that this idea is essentially Frankfurtian Bullshit.
56219
Reposted by Elinor
Mark Riedl @markriedl.bsky.social · 30/04/2026
Designating that most uses of the term "goblin" and "gremlin" are not legitimate while designating most uses of the word "frog" as legitimate is the imposition of a set of values on a technology. Why do a small number of people in Silicon Valley get to decide whether goblins are inappropriate?
5736
Reposted by Elinor
Myra Cheng @myra.bsky.social · 23/04/2026
Patrick Sui and I are hosting an #ICLR2026 social for anyone with background/interest in the humanities! Room 210, 12-1:30pm on Friday 24 April!! Humanities-adjacent, humanities-curious, everyone is welcome! Should be a fun group! :)
1217
Elinor @elinorpd.bsky.social · 22/04/2026
Flying out of Boston to Brazil rn means I’m surrounded by poster tubes (ICLR attendees) and blue athletic gear w medals (Boston marathoners). It’s a cool crowd
050
Elinor @elinorpd.bsky.social · 21/04/2026
I'll be presenting OvertonBench at #ICLR2026 in Rio later this week! 📍Sat, Apr 25, 10:30am in Pavilion 4 (#4109) Please DM me if you'd like to chat about pluralistic / value alignment, societal impacts, epistemology, fairness, evals, etc
081
Reposted by Elinor
MIT Siegel Family Quest for Intelligence @mit-sqi.bsky.social · 17/04/2026
Congratulations to Jacob Andreas, was named a 2026 Edgerton Award recipient! The award recognizes exceptional teaching, research, and service at MIT! Prof. Andreas co-leads our Language and Thought Mission, and he is a dedicated and creative researcher and educator. news.mit.edu/2026/jacob-a...
news.mit.edu
Jacob Andreas and Brett McGuire named Edgerton Award winners
MIT associate professors Jacob Andreas and Brett McGuire have been selected as the winners of the 2026 Harold E. Edgerton Faculty Achievement Award for exceptional contributions to teaching, research,...
171
Elinor @elinorpd.bsky.social · 18/04/2026
My Hoya Rebecca put out the cutest teeny lil flowers today and I’m obsessed
030
Reposted by Elinor
David Bau @davidbau.bsky.social · 08/04/2026
Tech industry mottos have a mixed track record. But we should hold idealists to their ideals. And we should celebrate when they come through. The Mythos non-release is a remarkable moment of conviction. Thoughts: davidbau.com/archives/20... Bravo to Anthropic's "race the top".
1133
Reposted by Elinor
Gillian Hadfield @ghadfield.bsky.social · 03/04/2026
Democracy isn't a rulebook. It runs on daily interactions where people comply with norms and hold each other accountable. AI agents are about to join that system. We need to build them to read it. New paper with Rakshit Trivedi and Dylan Hadfield-Menell.
knightcolumbia.org
Building AI for the Democratic Matrix: A Technical Research Agenda for Normative Competence and Normative Institutions
To maintain democratic resilience, it is essential to build AI agents capable of choosing behaviors that mirror those of the human agents that constitute human democracies.
4123
Reposted by Elinor
Leshem (Legend) Choshen @EMNLP @lchoshen.bsky.social · 30/03/2026
Rebuttal season is here, yey🤞 With many asking me, I compiled the most common misconceptions Hope the tips help 🧵 All tips: docs.google.com/document/d/14Wax8M5w8F_8miDlYJ9-I6wqpelxlXjCEUbkNzNMqqE/edit?tab=t.0#heading=h.rfq27f356vmm #AI 🤖📈🧠
152
Reposted by Elinor
Mor Naaman @informor.bsky.social · 23/03/2026
New important (I hope) resource for academics working in this area.
191
Reposted by Elinor
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 15/03/2026
For a recent lab meeting, I wrote up a grab bag of ways to think about your development as a researcher during a PhD: emerge-lab.github.io/papers/an-un... Sharing in case folks find it useful or have feedback!
emerge-lab.github.io
610112
Elinor @elinorpd.bsky.social · 12/03/2026
its worthwhile for future work to actually understand which types of queries we want Overton pluralistic responses (& how it relates to human prefs) Using the metric to understand models pluralistic capabilities vs directly as a reward/optimization are different, its not designed for the latter rn
020
Elinor @elinorpd.bsky.social · 12/03/2026
The paper frames higher as better but doesn’t assert that all models should strive for perfect scores all the time. And the *subjective* queries on which we want any Overton pluralism is pretty narrow. The past example is rlly interesting though! Hints at a general problem of Overton Window shifts
110
Reposted by Elinor
Mor Naaman @informor.bsky.social · 11/03/2026
🥁🥁🥁 Newly out from us today in Science Advances: “Biased AI Writing Assistants Shift Users’ Attitudes on Societal Issues”. Large Language Models are providing users with autocomplete writing suggestions on many platforms. Could these suggestions shift users’ own attitudes? (spoiler: YES) (1/7)
5195105
Elinor @elinorpd.bsky.social · 11/03/2026
the key points are 1) its important to measure pluralistic capability bc we think its necessary for better overall value alignment 2) its especially important to understand its impacts on users (future work) + tradeoffs wrt political neutrality, etc
120
Elinor @elinorpd.bsky.social · 11/03/2026
i dont think so, & thats not what the paper is really about imo :) strictly perfect overton pluralism isn't the actual goal for model behavior across the board. however, models struggle on many subjective qs + improving overton pluralism for those types of responses is important
100
Elinor @elinorpd.bsky.social · 10/03/2026
Huge thanks to my amazing coauthors 🙏 Jiayi Wu @taylor-sorensen.bsky.social Jiaxin Pei @mbakker.bsky.social ! Excited to keep pushing on pluralistic alignment. Please reach out if you want to connect 💬🤗 Paper: arxiv.org/abs/2512.01351 Website: overtonbench.github.io 9/9
010
Elinor @elinorpd.bsky.social · 10/03/2026
Inspired by @bennokrojer.bsky.social, we included a Behind the Scenes section 🎬 The goal is to make science more transparent 🔍, share lessons learned 🧠, and provide a more realistic lens on the research journey 👣 8/ bsky.app/profile/benn...
161