Reposted by ElinorLeshem (Legend) Choshen @EMNLP @lchoshen.bsky.social · 04/09/2026She's a sexy fox 🦊 היא חתולה סקסית 😻 People believe translation is solved, because we dropped the bar. But the above is something no translation system (aka LLM) is close to doing. We made worldwide challenges and keep collecting so you can read use it join the authors... 181
Reposted by ElinorIsabelle Augenstein @iaugenstein.bsky.social · 03/09/2026Our paper on measuring epistemic diversity in LLMs is accepted to #EMNLP2026! We find that diversity has improved, though LLM output is much less diverse than Web search, especially for non-English. Paper: arxiv.org/pdf/2510.04226 Code/data: github.com/dwright37/ll... #NLProc @copenlu.bsky.social 1263
Reposted by ElinorShomir Wilson @shomir.bsky.social · 24/08/2026I'm pleased to share a preprint of our accepted EMNLP main conference paper "Affective Context Amplifies Sycophancy in LLM Responses". Sycophancy, an unnatural level of agreement or praise for a person's ideas, is a problem for chatbots, sometimes with disastrous outcomes. arxiv.org/abs/2608.212...arxiv.orgAffective Context Amplifies Sycophancy in LLM ResponsesAs conversational companions, large language models (LLMs) often have access to users' emotional states. We study how this affective context modulates LLM sycophancy in subjective, evaluative interact... 0141
Reposted by ElinorLucy Li @lucy3.bsky.social · 07/08/2026Spent a good amount of this summer digging through a 20+ person team of teachers' free-text annotations, learning what "scaffolding" and "push for rigor" means, and iterating on this pipeline. We find that AI tutors, by default, frequently over-scaffold and rarely push for rigor. 27412
Reposted by ElinorHenry Yuen @henryyuen.bsky.social · 01/08/2026I spent many hours, days, nights, weekends in cafes, in my office, at home, trying to understand Ran Raz's classical parallel repetition theorem and whether I could prove a quantum version of it. This period of struggle was important for me. 1835
Elinor @elinorpd.bsky.social · 23/07/2026Pluralistic alignment is thriving as a research agenda yet failing at its goal: making the AI systems people actually use more pluralistic🌈 🚨New position paper: we argue adoption in deployed models should be the fields main goal & we provide a roadmap of how to get there 🧵1/ 3212
Reposted by ElinorDavid Jurgens @davidjurgens.bsky.social · 13/07/2026With my ACL PC abilities, I wanted to characterize who these folks are, whether they're actually "outsiders", whether their papers are much less likely to be accepted, is this growth just LLM slop? The whole analsis is here medium.com/@jurgens_245...medium.comIs the ACL Rolling Review actually broken?ARR is broken! ARR is being overwhelmed with papers from outsiders! LLM-generated slop is ruining our peer review! There are not enough… 12811
Reposted by ElinorJonathan Stray @jonathanstray.com · 23/10/2025GreenEarth is creating open source AI-driven recommender infrastructure for BlueSky. Type a prompt, see your feed change. We are here for the users, the builders, the dreamers. Join us. greenearthsocial.substack.com/p/introducin...greenearthsocial.substack.comIntroducing GreenEarthWe're building advanced open source algorithms for social media 1010427
Reposted by ElinoriNaturalist @inaturalist.bsky.social · 06/07/2026We're looking for a CV/ML Engineer to help us improve the machine learning systems that power iNaturalist's species identification and geographic range modeling. If you're excited to help build tools that help millions of people engage with nature, we'd love to hear from you! Apply: buff.ly/YZqaW6c 03921
Reposted by ElinorEhud Reiter @ehudreiter.bsky.social · 08/06/2026New blog: I am worried by NLP research culture NLG and NLP are mostly much better in 2026 than when I got my PhD in 1990. Unfortunately research culture has gotten *worse” in this period, which really worries me as I retire. ehudreiter.com/2026/06/08/n...ehudreiter.comI am worried by NLP research cultureIn most ways NLG and NLP are much better in 2026 than when I got my PhD in 1990. Unfortunately research culture has gotten *worse” in this period, which really worries me as I retire. We have… 1144
Elinor @elinorpd.bsky.social · 06/06/2026excited to share that i'll be pursuing my phd in computer science at @csail.mit.edu starting this fall 🥳 🎓 i'm so grateful to be coadvised by the literal dream team: Jacob Andreas, @mbakker.bsky.social, and @mitchellgordon.bsky.social 🙌 4240
Reposted by ElinorBenno Krojer @bennokrojer.bsky.social · 05/06/2026My first last-author paper is out! If you saw this dog below and someone showed you the second image, would you consider them the same word/concept? (more examples in Ada's thread) We study if VLMs agree with humans on this and revisit old questions around shape vs. texture bias in vision 1124
Reposted by ElinorJessica Hullman @jessicahullman.bsky.social · 28/05/2026Thoughts on metascientific consequences of AI-generated slides & ideas diluting the impression that speakers are commited to what they present. Science runs on personal attachment more than we admit. If it were a cake mix, how wouldn we add back an egg? statmodeling.stat.columbia.edu/2026/05/28/w...statmodeling.stat.columbia.edu What if scientists really were dispassionate observers, communicating ideas without irrational commitment? Look here, says AI. | Statistical Modeling, Causal Inference, and Social Science 1296
Reposted by ElinorIlias Chalkidis @kiddothe2b.bsky.social · 06/05/2026📢 Our paper 🤯🧠 "Brainrot: Deskilling and Addiction are Overlooked AI Risks" has been accepted at the ACM Fairness, Accountability & Transparency (FAccT) conference 2026. The preprint is available: 👉 arxiv.org/abs/2605.03512 TL;DR 🧵 follows 👇 1/5arxiv.orgBrainrot: Deskilling and Addiction are Overlooked AI RisksThe scope of AI safety and alignment work in generative artificial intelligence (GenAI) has so far mostly been limited to harms related to: (a) discrimination and hate speech, (b) harmful/inappropriat... 1215
Reposted by ElinorJoachim Baumann @joachimbaumann.bsky.social · 01/05/2026Can you boost your AI review scores by asking an LLM to rewrite your paper? Yes! We call it paper laundering Our @icmlconf.bsky.social spotlight paper argues current AI reviewers aren't ready to automate peer review, and outlines what a science of peer review automation should look like 🧵👇 #ICML2026 44314
Reposted by ElinorPetter Törnberg @pettertornberg.com · 01/05/2026LLMs have been widely reported as left-wing biased. The finding has shaped policy and debate — with Trump banning "Woke AI". Our new paper challenges this story. It's not that the models are biased. It's that they think the auditor is. 🧵 w/ michelleschimmel.bsky.socialarxiv.org/pdf/2604.27633 415750
Reposted by ElinorShahan Ali Memon @shahanmemon.bsky.social · 30/04/2026Our position paper, "AI Welfare Is Bullshit" just got accepted to ICML 2026 @icmlconf.bsky.social! The AI welfare agenda has already begun to attract institutional investment at organizations such an @anthropic.com. We argue that this idea is essentially Frankfurtian Bullshit. 56219
Reposted by ElinorMark Riedl @markriedl.bsky.social · 30/04/2026Designating that most uses of the term "goblin" and "gremlin" are not legitimate while designating most uses of the word "frog" as legitimate is the imposition of a set of values on a technology. Why do a small number of people in Silicon Valley get to decide whether goblins are inappropriate? 5736
Reposted by ElinorMyra Cheng @myra.bsky.social · 23/04/2026Patrick Sui and I are hosting an #ICLR2026 social for anyone with background/interest in the humanities! Room 210, 12-1:30pm on Friday 24 April!! Humanities-adjacent, humanities-curious, everyone is welcome! Should be a fun group! :) 1217
Elinor @elinorpd.bsky.social · 22/04/2026Flying out of Boston to Brazil rn means I’m surrounded by poster tubes (ICLR attendees) and blue athletic gear w medals (Boston marathoners). It’s a cool crowd 050
Elinor @elinorpd.bsky.social · 21/04/2026I'll be presenting OvertonBench at #ICLR2026 in Rio later this week! 📍Sat, Apr 25, 10:30am in Pavilion 4 (#4109) Please DM me if you'd like to chat about pluralistic / value alignment, societal impacts, epistemology, fairness, evals, etc 081
Reposted by ElinorMIT Siegel Family Quest for Intelligence @mit-sqi.bsky.social · 17/04/2026Congratulations to Jacob Andreas, was named a 2026 Edgerton Award recipient! The award recognizes exceptional teaching, research, and service at MIT! Prof. Andreas co-leads our Language and Thought Mission, and he is a dedicated and creative researcher and educator. news.mit.edu/2026/jacob-a...news.mit.eduJacob Andreas and Brett McGuire named Edgerton Award winnersMIT associate professors Jacob Andreas and Brett McGuire have been selected as the winners of the 2026 Harold E. Edgerton Faculty Achievement Award for exceptional contributions to teaching, research,... 171
Elinor @elinorpd.bsky.social · 18/04/2026My Hoya Rebecca put out the cutest teeny lil flowers today and I’m obsessed 030
Reposted by ElinorDavid Bau @davidbau.bsky.social · 08/04/2026Tech industry mottos have a mixed track record. But we should hold idealists to their ideals. And we should celebrate when they come through. The Mythos non-release is a remarkable moment of conviction. Thoughts: davidbau.com/archives/20... Bravo to Anthropic's "race the top". 1133
Reposted by ElinorGillian Hadfield @ghadfield.bsky.social · 03/04/2026Democracy isn't a rulebook. It runs on daily interactions where people comply with norms and hold each other accountable. AI agents are about to join that system. We need to build them to read it. New paper with Rakshit Trivedi and Dylan Hadfield-Menell.knightcolumbia.orgBuilding AI for the Democratic Matrix: A Technical Research Agenda for Normative Competence and Normative InstitutionsTo maintain democratic resilience, it is essential to build AI agents capable of choosing behaviors that mirror those of the human agents that constitute human democracies. 4112
Reposted by ElinorLeshem (Legend) Choshen @EMNLP @lchoshen.bsky.social · 30/03/2026Rebuttal season is here, yey🤞 With many asking me, I compiled the most common misconceptions Hope the tips help 🧵 All tips: docs.google.com/document/d/14Wax8M5w8F_8miDlYJ9-I6wqpelxlXjCEUbkNzNMqqE/edit?tab=t.0#heading=h.rfq27f356vmm #AI 🤖📈🧠 152
Reposted by ElinorMor Naaman @informor.bsky.social · 23/03/2026New important (I hope) resource for academics working in this area. 191
Reposted by ElinorEugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 15/03/2026For a recent lab meeting, I wrote up a grab bag of ways to think about your development as a researcher during a PhD: emerge-lab.github.io/papers/an-un... Sharing in case folks find it useful or have feedback!emerge-lab.github.io 610112
Reposted by ElinorMor Naaman @informor.bsky.social · 11/03/2026🥁🥁🥁 Newly out from us today in Science Advances: “Biased AI Writing Assistants Shift Users’ Attitudes on Societal Issues”. Large Language Models are providing users with autocomplete writing suggestions on many platforms. Could these suggestions shift users’ own attitudes? (spoiler: YES) (1/7) 5195105
Elinor @elinorpd.bsky.social · 10/03/2026There's been a lot of excitement about pluralistic value alignment 🌈 — AI that reflects the full range of human perspectives But no formal way to benchmark whether we're actually making progress. 🤔 Introducing 𝐎𝐕𝐄𝐑𝐓𝐎𝐍𝐁𝐄𝐍𝐂𝐇. 🎉Accepted to #ICLR2026 1/n 🧵 1231
Reposted by ElinorLeshem (Legend) Choshen @EMNLP @lchoshen.bsky.social · 09/03/2026Do LLMs Benefit from Their Own Words?🤔 In multi-turn chats, models are typically given their own past responses as context. But do their own words always help… Or are they more often a waste of compute and a distraction? 🧵 arxiv.org/abs/2602.24287 2374
Reposted by ElinorLucy Li @lucy3.bsky.social · 03/03/2026Models are now expert math solvers, and so AI for math education is receiving increasing attention. Our new preprint evaluates 11 VLMs on our QA benchmark, DrawEduMath. We highlight a startling gap: models perform less well on inputs from K-12 students who need more help. 🧵 43512
Reposted by ElinorAlexandra Olteanu @aolteanu.bsky.social · 03/03/2026Yesterday was my last day at MSR. We recently learned that our roles were eliminated, and with them our little FATE Montreal team. I joined MSR a bit over 7.5 years ago while on active chemotherapy, and being at MSR has overlapped with so much change in my life. 4346
Reposted by ElinorEmma Pierson @emmapierson.bsky.social · 26/02/2026Our paper, "What's in My Human Feedback", received an oral presentation at ICLR! Our method automatically+interpretably identifies preferences in human feedback data; we use this to improve personalization + safety. Reach out if you have data/use cases to apply this to! arxiv.org/pdf/2510.26202 0283
Reposted by ElinorBenno Krojer @bennokrojer.bsky.social · 11/02/2026Finally we do test it empirically: finding some models where the embedding matrix of the LLM already provides decently interpretable nearest neighbors But this was not the full story yet... @mariusmosbach.bsky.social and @elinorpd.bsky.social nudged me to use contextual embeddings 111
Elinor @elinorpd.bsky.social · 11/02/2026Really cool new work with surprising results! Highly recommend checking out the demo 👀 030
Reposted by ElinorDavid Rand @dgrand.bsky.social · 04/02/2026Grok fact-checks our paper on Grok fact-checking - and it approves! 1297
Reposted by ElinorShaily @shaily99.bsky.social · 02/02/2026🎭 How do LLMs (mis)represent culture? 🧮 How often? 🧠 Misrepresentations = missing knowledge? spoiler: NO! At #CHI2026 we are bringing ✨TALES✨ a participatory evaluation of cultural (mis)reps & knowledge in multilingual LLM-stories for India 📜 arxiv.org/abs/2511.21322 1/10 14722
Elinor @elinorpd.bsky.social · 30/01/2026Potato is a great platform for researchers! Highly recommend (plus a great development team behind it) 010
Reposted by ElinorHanna Wallach @hannawallach.bsky.social · 29/01/2026Microsoft Research NYC is hiring a researcher in the space of AI and society! 26240
Reposted by ElinorDavid Bau @davidbau.bsky.social · 26/01/2026What should academics be doing right now? I have been writing up some thoughts on what the research says about effective action, and what universities specifically can do. davidbau.github.io/poetsandnurs... It's on GitHub. Suggestions and pull requests welcome. github.com/davidbau/poe... 03715
Elinor @elinorpd.bsky.social · 23/01/2026I'll be presenting this work on January 25th (Hall 2, poster 41) at #AAAI2026 in Singapore! Please stop by and reach out if you'd like to chat 😁 030
Elinor @elinorpd.bsky.social · 23/01/2026🎉 Excited to share our new paper which was accepted to #AAAI2026! As LLMs become increasingly used as sources of factual knowledge, we ask: Do they perform equitably across users of different backgrounds? 🧵⬇️ 1/6 121
Reposted by ElinorAngelina Wang @angelinawang.bsky.social · 12/12/2025Most LLM evals use API calls or offline inference, testing models in a memory-less silo. Our new Patterns paper shows this misses how LLMs actually behave in real user interfaces, where personalization and interaction history shape responses: arxiv.org/abs/2509.19364 13811
Reposted by ElinorarXiv cs.AI Artificial Intelligence @csai-bot.bsky.social · 02/12/2025Elinor Poole-Dayan, Jiayi Wu, Taylor Sorensen, Jiaxin Pei, Michiel A. Bakker: Benchmarking Overton Pluralism in LLMs arxiv.org/abs/2512.01351 arxiv.org/pdf/2512.01351 arxiv.org/html/2512.01351 021
Reposted by ElinorGautam Kamath @gautamkamath.com · 19/11/2025Thoughtful (as always) blog post from Nicholas Carlini. "Are large language models worth it?" A nice read giving his perspective on risks of ML models. Post: nicholas.carlini.com/writing/2025... For people who prefer, this is the video of the talk from @colmweb.org www.youtube.com/watch?v=PngH... 13411
Reposted by ElinorAvijit Ghosh @evijit.io · 13/11/2025Extremely thrilled to talk about our new paper: "Who Evaluates AI’s Social Impacts? Mapping Coverage And Gaps In First And Third Party Evaluations". This is the first big project output from the @eval-eval.bsky.social coalition! Thread below: 1207