Sign in

Myra Cheng

@myra.bsky.social
2.9K followers 147 following 46 posts

PhD candidate @ Stanford NLP myracheng.github.io

PostsRepliesMedia
Reposted by Myra Cheng
Melanie Walsh @mellymeldubs.bsky.social · 24/06/2026
Excited to share this. @neel2112.bsky.social, @mariaa.bsky.social, and I analyzed 500K anonymous ChatGPT convos (shared w/ consent from WildChat) to see if people were generating fiction. We found tons of stories, fanfiction & erotica. Many users iterated on the same stories for days and weeks.
Screenshot of paper abstract that reads: 

AI FICTION IN THE WILD Neel Gupta  Maria Antoniak  Melanie Walsh

Some professional authors are beginning to use AI tools to help produce their fiction writing. Are readers using AI to generate fiction, too? Drawing on over 500,000 anonymized, English-language ChatGPT-user conversations (Zhao et al.), we find that more than one third of the conversations involve some form of fiction generation—including original stories, roleplay, fanfiction, and erotica. This AI-generated fiction is notably dominated by power users. We identify common fiction generation patterns and profiles among these users, including what we call infinite story demanders, who repeatedly request and revise variations of the same or similar narratives over extended periods of time. We show that users especially gravitate toward fanfiction and erotica, and that they are broadly drawn to generic forms, repetition, immediacy, and niche combinations of story elements. Our findings motivate two theoretical provocations. First, we argue that AI technologies may lead to a shift in the conventional relationship between the author and reader, potentially producing what we call a solipsistic reader-writer, who both generates and consumes fiction within a closed conversational loop, interacting with a machine rather than a human other. Second, we note that LLMs enable interactivity, play, and permutation in ways that are seemingly pleasurable for users, raising questions about where AI will fit into contemporary storytelling and entertainment ecosystems. We situate these developments within broader transformations in literature and media, including self-publishing, fanfiction, and pornography, and suggest that AI-generated fiction shares structural affinities with on-demand, personalized, and repetitive cultural forms.
515645
Reposted by Myra Cheng
Naomi Saphra @nsaphra.bsky.social · 15/06/2026
We don’t always know what problems are hard for LLMs. So devs evaluate on tasks HUMANS find hard or on broad benchmarks. What if we could instead anticipate which scenarios a model will fail on—all without evaluating specific input examples? 🧵NEW PAPER by @jenniferlumeng.bsky.social
313734
Myra Cheng @myra.bsky.social · 23/04/2026
In case it's of interest to your circles - @richardjeanso.bsky.social @mariaa.bsky.social @naitian.org @lucy3.bsky.social
040
Myra Cheng @myra.bsky.social · 23/04/2026
Formal pitch: Humanities has much to offer AI/NLP in methods & theory for understanding language, culture, narrative, subjectivity, mind, emotion, etc. Topics: co-intelligence, co-creativity, narrative understanding, cultural analytics, literary NLP, AI literacy, cognition, and lots more!!
2100
Myra Cheng @myra.bsky.social · 23/04/2026
Patrick Sui and I are hosting an #ICLR2026 social for anyone with background/interest in the humanities! Room 210, 12-1:30pm on Friday 24 April!! Humanities-adjacent, humanities-curious, everyone is welcome! Should be a fun group! :)
1217
Myra Cheng @myra.bsky.social · 13/04/2026
happy to chat about sycophancy, public perceptions of AI, AI interaction through a pragmatic lens, anthropomorphism, or anything else :-) please reach out! Assumptions preprint: arxiv.org/pdf/2604.03058
arxiv.org
175
Myra Cheng @myra.bsky.social · 13/04/2026
In Barcelona for #chi2026! Presenting our work on eliciting LLMs' assumptions about users, and how this mismatches with user expectations, in the Tues poster session! (Spoiler: users assume that LLMs give objective info much more than they actually do, which leads to sycophancy)
1181
Reposted by Myra Cheng
M.J. Crockett @mjcrockett.bsky.social · 22/02/2026
Many are appropriately outraged by Altman’s comments here implying that raising a human child is akin to “training” an AI model. This is part of a broader pattern where AI industry leaders use language that collapses the boundary between human and machine. 🧵/
28495193
Reposted by Myra Cheng
Kaitlyn Zhou @kaitlynzhou.bsky.social · 06/11/2025
No better time to start learning about that #AI thing everyone's talking about... 📢 I'm recruiting PhD students in Computer Science or Information Science @cornellbowers.bsky.social! If you're interested, apply to either department (yes, either program!) and list me as a potential advisor!
Photo of Cornelll University building surrounded by colorful trees
2259
Reposted by Myra Cheng
Kaitlyn Zhou @kaitlynzhou.bsky.social · 21/10/2025
As of June 2025, 66% of Americans have never used ChatGPT. Our new position paper, Attention to Non-Adopters, explores why this matters: AI research is being shaped around adopters—leaving non-adopters’ needs, and key LLM research opportunities, behind. arxiv.org/abs/2510.15951
A circular flow diagram that compares current and proposed practices for LLM development using data from adopters and non-adopters. Three gray boxes represent current practices: “R&D,” “Chat Models,” and “Adopters’ Needs and Usage Data,” connected in a clockwise loop with black arrows. A blue box labeled “Non-adopters’ Needs and Usage Data” adds a proposed feedback path, shown with blue arrows, linking non-adopter data back to R&D and adopters’ data.
23912
Reposted by Myra Cheng
Kaitlyn Zhou @kaitlynzhou.bsky.social · 03/10/2025
I'll be at COLM next week! Let me know if you want to chat! @colmweb.org @neilrathi.bsky.social will be presenting our work on multilingual overconfidence in language models and the effects on human overreliance! arxiv.org/pdf/2507.06306
arxiv.org
071
Reposted by Myra Cheng
Steve Rathje @steverathje.bsky.social · 01/10/2025
🚨 New preprint 🚨 Across 3 experiments (n = 3,285), we found that interacting with sycophantic (or overly agreeable) AI chatbots entrenched attitudes and led to inflated self-perceptions. Yet, people preferred sycophantic chatbots and viewed them as unbiased! osf.io/preprints/ps... Thread 🧵
Abstract and results summary
6193102
Myra Cheng @myra.bsky.social · 03/10/2025
Was a blast working on this with @cinoolee.bsky.social @pranavkhadpe.bsky.social, Sunny Yu, Dyllan Han, and @jurafsky.bsky.social !!! So lucky to work with this wonderful interdisciplinary team!!💖✨
1223
Myra Cheng @myra.bsky.social · 03/10/2025
While our work focuses on interpersonal advice-seeking, concurrent work by @steverathje.bsky.social @jayvanbavel.bsky.social et al. finds similar patterns for political topics, where sycophantic AI also led to more extreme attitudes when users discussed gun control, healthcare, immigration, etc.!
1357
Myra Cheng @myra.bsky.social · 03/10/2025
There is currently little incentive for developers to reduce sycophancy. Our work is a call to action: we need to learn from the social media era and actively consider long-term wellbeing in AI development and deployment. Read our preprint: arxiv.org/pdf/2510.01395
arxiv.org
35514
Myra Cheng @myra.bsky.social · 03/10/2025
Despite sycophantic AI’s reduction of prosocial intentions, people also preferred it and trusted it more. This reveals a tension: AI is rewarded for telling us what we want to hear (immediate user satisfaction), even when it may harm our relationships.
Rightness judgment is higher and repair likelihood is lower for sycophantic AIResponse quality, return likelihood, and trust are higher for sycophantic AI
16818
Myra Cheng @myra.bsky.social · 03/10/2025
Next, we tested the effects of sycophancy. We find that even a single interaction with sycophantic AI increased users’ conviction that they were right and reduced their willingness to apologize. This held both in controlled, hypothetical vignettes and live conversations about real conflicts.
Description of Study 2 (hypothetical vignettes) and Study 3 (live interaction) where self-attributed wrongness and desire to initiate repair decrease, while response quality and trust increases.
16414
Myra Cheng @myra.bsky.social · 03/10/2025
We focus on the prevalence and harms of one dimension of sycophancy: AI models endorsing users’ behaviors. Across 11 AI models, AI affirms users’ actions about 50% more than humans do, including when users describe harmful behaviors like deception or manipulation.
Description of Study 1, where we characterize the prevalence of social sycophancy and find it to be highly prevalent across leading AI models
17616
Myra Cheng @myra.bsky.social · 03/10/2025
AI always calling your ideas “fantastic” can feel inauthentic, but what are sycophancy’s deeper harms? We find that in the common use case of seeking AI advice on interpersonal situations—specifically conflicts—sycophancy makes people feel more right & less willing to apologize.
Screenshot of paper title: Sycophantic AI Decreases Prosocial Intentions and Promotes Dependence
5544270
Myra Cheng @myra.bsky.social · 05/08/2025
Thoughtful NPR piece about ChatGPT relationship advice! Thanks for mentioning our research :)
0150
Myra Cheng @myra.bsky.social · 04/08/2025
Congrats Maria!! All the best!!
030
Reposted by Myra Cheng
Alexandra Olteanu @aolteanu.bsky.social · 29/07/2025
#acl2025 I think there is plenty of evidence for the risks of anthropomorphic AI behavior and design (re: keynote) -- find @myra.bsky.social and I if you want to chat more about this or our "Dehumanizing Machines" ACL 2025 paper
0111
Reposted by Myra Cheng
Dr Abeba Birhane @abeba.blacksky.app · 25/06/2025
New paper hot off the press www.nature.com/articles/s41... We analysed over 40,000 computer vision papers from CVPR (the longest standing CV conf) & associated patents tracing pathways from research to application. We found that 90% of papers & 86% of downstream patents power surveillance 1/
nature.com
Computer-vision research powers surveillance technology - Nature
An analysis of research papers and citing patents indicates the extensive ties between computer-vision research and surveillance.
341010554
Myra Cheng @myra.bsky.social · 28/06/2025
Aw thanks!! :)
010
Myra Cheng @myra.bsky.social · 12/06/2025
Paper: arxiv.org/pdf/2502.13259 Code: github.com/myracheng/hu... Thanks to my wonderful collaborators Sunny Yu and @jurafsky.bsky.social and everyone who helped along the way!!
arxiv.org
200
Myra Cheng @myra.bsky.social · 12/06/2025
So we built DumT, a method using DPO + HumT to steer models to be less human-like without hurting performance. Annotators preferred DumT outputs for being: 1) more informative and less wordy (no extra “Happy to help!”) 2) less deceptive and more authentic to LLMs’ capabilities.
Plots showing that DumT reduces MeanHumT and has higher performance on RewardBench than the baseline models.
120
Myra Cheng @myra.bsky.social · 12/06/2025
We also develop metrics for implicit social perceptions in language, and find that human-like LLM outputs correlate with perceptions linked to harms: warmth and closeness (→ overreliance), and low status and femininity (→ harmful stereotypes).
human-like LLM outputs are strongly positively correlated with social closeness, femininity, and warmth (r = 0.87, 0.47, 0.45), and strongly negatively correlated with status (r = 0.80).
210
Myra Cheng @myra.bsky.social · 12/06/2025
First, we introduce HumT (Human-like Tone), a metric for how human-like a text is, based on relative LM probabilities. Measuring HumT across 5 preference datasets, we find that preferred outputs are consistently less human-like.
bar plot showing that human-likeness is lower in preferred responses
131
Myra Cheng @myra.bsky.social · 12/06/2025
Do people actually like human-like LLMs? In our #ACL2025 paper HumT DumT, we find a kind of uncanny valley effect: users dislike LLM outputs that are *too human-like*. We thus develop methods to reduce human-likeness without sacrificing performance.
Screenshot of first page of the paper HumT DumT: Measuring and controlling human-like language in LLMs
1236
Myra Cheng @myra.bsky.social · 22/05/2025
thanks!! looking forward to seeing your submission as well :D
010
Myra Cheng @myra.bsky.social · 22/05/2025
thanks Rob!!
000
Myra Cheng @myra.bsky.social · 21/05/2025
We also apply ELEPHANT to identify sources of sycophancy (in preference datasets) and explore mitigations. Our work enables measuring social sycophancy to prevent harms before they happen. Preprint: arxiv.org/abs/2505.13995 Code: github.com/myracheng/el...
github.com
GitHub - myracheng/elephant
Contribute to myracheng/elephant development by creating an account on GitHub.
030
Myra Cheng @myra.bsky.social · 21/05/2025
Oops, yes! arxiv.org/abs/2505.13995
arxiv.org
Social Sycophancy: A Broader Understanding of LLM Sycophancy
A serious risk to the safety and utility of LLMs is sycophancy, i.e., excessive agreement with and flattery of the user. Yet existing work focuses on only one aspect of sycophancy: agreement with user...
030
Myra Cheng @myra.bsky.social · 21/05/2025
Grateful to work with Sunny Yu (undergrad!!!) @cinoolee.bsky.social @pranavkhadpe.bsky.social @lujain.bsky.social @jurafsky.bsky.social on this! Lots of great cross-disciplinary insights:)
170
Myra Cheng @myra.bsky.social · 21/05/2025
We also apply ELEPHANT to identify sources of sycophancy (in preference datasets) and explore mitigations. Our work enables measuring social sycophancy to prevent harms before they happen. Preprint: arxiv.org/abs/2505.13995 Code: github.com/myracheng/el...
020
Myra Cheng @myra.bsky.social · 21/05/2025
We apply ELEPHANT to 8 LLMs across two personal advice datasets (Open-ended Questions & r/AITA). LLMs preserve face 47% more than humans, and on r/AITA, LLMs endorse the user’s actions in 42% of cases where humans do not.
261
Myra Cheng @myra.bsky.social · 21/05/2025
By defining social sycophancy as excessive preservation of the user’s face (i.e., their desired self-image), we capture sycophancy in these complex, real-world cases. ELEPHANT, our evaluation framework, detects 5 face-preserving behaviors.
3161
Myra Cheng @myra.bsky.social · 21/05/2025
Prior work only looks at whether models agree with users’ explicit statements vs. a ground truth. But for real-world queries, which often contain implicit beliefs and do not have ground truth, sycophancy can be subtler and more dangerous.
1111
Myra Cheng @myra.bsky.social · 21/05/2025
Dear ChatGPT, Am I the Asshole? While Reddit users might say yes, your favorite LLM probably won’t. We present Social Sycophancy: a new way to understand and measure sycophancy as how LLMs overly preserve users' self-image.
614735
Myra Cheng @myra.bsky.social · 02/05/2025
super interesting!!
010
Myra Cheng @myra.bsky.social · 02/05/2025
Thanks!! :D
000
Myra Cheng @myra.bsky.social · 02/05/2025
Metaphors offer a powerful lens into how people understand AI and how we can design and build AI more responsibly! Check out the full paper: arxiv.org/pdf/2501.18045
arxiv.org
130
Myra Cheng @myra.bsky.social · 02/05/2025
Metaphors shape trust and adoption. Seeing AI as a "friend" or "teacher" is associated with higher trust, while "thief" signals the opposite. We find that women, older adults, and POC anthropomorphize AI more; and, the latter two groups have higher trust and adoption.
140
Myra Cheng @myra.bsky.social · 02/05/2025
We find that: ➕ Anthropomorphism is rising over time: people are seeing AI as more human-like and agentic. ➕ Warmth is rising over time.
121
Myra Cheng @myra.bsky.social · 02/05/2025
To quantify these perceptions, we combine AnthroScore (anthroscore.stanford.edu) and a computational SCM model (aclanthology.org/2021.acl-lon...) to measure anthropomorphism, warmth, and competence from the metaphors.
130
Myra Cheng @myra.bsky.social · 02/05/2025
We identify 20 dominant metaphors—ranging from "friend" to "god" to "thief"—and how their prevalence is shifting over time: warm, human-like metaphors (friend, assistant) are rising while mechanical ones that signal competence (computer, search engine) are declining.
1101
Myra Cheng @myra.bsky.social · 02/05/2025
How does the public conceptualize AI? Rather than self-reported measures, we use metaphors to understand the nuance and complexity of people’s mental models. In our #FAccT2025 paper, we analyzed 12,000 metaphors collected over 12 months to track shifts in public perceptions.
34914
Myra Cheng @myra.bsky.social · 27/04/2025
This draws on work from last summer at MSR MTL, with @aolteanu.bsky.social, Su Lin Blodgett, @uhleeeeeeeshuh.bsky.social, and @lisaegede.bsky.social! Check out the full post: iclr-blogposts.github.io/2025/blog/an...
iclr-blogposts.github.io
“I Am the One and Only, Your Cyber BFF”: Understanding the Impact of GenAI Requires Understanding the Impact of Anthropomorphic AI | ICLR Blogposts 2025
State-of-the-art generative AI (GenAI) systems are increasingly prone to anthropomorphic behaviors, i.e., to generating outputs that are perceived to be human-like. While this has led to scholars incr...
062
Myra Cheng @myra.bsky.social · 27/04/2025
In our blogpost, we outline key directions to provide scaffolding for this work: conceptual clarity (what counts as anthropomorphic behavior?), better terminology, mitigation strategies, and unpacking what practices lead to anthropomorphic AI.
120
Myra Cheng @myra.bsky.social · 27/04/2025
Despite growing concerns about risks such as anthropomorphic deception, emotional dependence, and overreliance, there’s still little systematic study of what exacerbates these risks or how to address them.
110