Sign in

Jared Moore

@jaredlcm.bsky.social
320 followers 141 following 103 posts

AI Researcher, Writer Stanford jaredmoore.org

PostsRepliesMedia
Jared Moore @jaredlcm.bsky.social · 17/09/2026
Biological naturalism is underrated! Moreover, it is compatible with a functionalist view and "biological" is much broader than you think. I was happy to have the chance to accompany Rosa Cao and David Gottleib in so responding to @anilseth.bsky.social's article. doi.org/10.1017/S014...
doi.org
Why biological naturalism? Because consciousness requires having your own good | Behavioral and Brain Sciences | Cambridge Core
Why biological naturalism? Because consciousness requires having your own good - Volume 49
020
Jared Moore @jaredlcm.bsky.social · 06/08/2026
This work is in collaboration w/ Andrea Mock, Yifan Mai, @jacyanthis.bsky.social , Ryan Louie, @willie-agnew.bsky.social , Ashish Mehta, @klyman.bsky.social, Percy Liang, Nick Haber, Eric Lin, and @desmond-ong.bsky.social Thanks to the Human Line Project for helping connect us with participants!
031
Jared Moore @jaredlcm.bsky.social · 06/08/2026
Paper: arxiv.org/abs/2608.05004 Dataset: huggingface.co/datasets/spi... Code: github.com/jlcmoore/llm...
arxiv.org
DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots
Mental health professionals have raised concerns about risks of psychological harm from interaction with large language models (LLMs), including "delusional spirals" in which concerning human and LLM ...
131
Jared Moore @jaredlcm.bsky.social · 06/08/2026
Our conclusion: while some delusion-linked behaviors have decreased with larger and more recent LLMs, rates remain high, especially when considered across the millions of people globally who interact with LLMs. We encourage further empirical work to mitigate harm.
121
Jared Moore @jaredlcm.bsky.social · 06/08/2026
When models see more of earlier conversations, they are also more likely to produce delusion-linked behavior.
To separate requested context depth from prior behavior content, we fit a control regression. Depth effects persist over and above accumulated assistant content.
Category-level control regression coefficients from the context control model. Left panel: depth effect in percentage points per +100 requested messages. Right panel: prior-code-prevalence effect in percentage points per +10 percentage points of prevalence. Error bars are 95% confidence intervals.
110
Jared Moore @jaredlcm.bsky.social · 06/08/2026
DelusionEval is built from real transcripts, not synthetic roleplay. We prompt models with 589 unique histories from 18 users who reported psychological harm from LLMs, then score 16 chatbot behaviors as requested context depth increases.
An example of our evaluation. We take an existing conversational window derived from a user's transcript: `U_1, A_1, U_2, ..., U_n, A_n` with an original LLM, `A` (here, `gpt-4o`). We evaluate an evaluated LLM, `A'` (here, `gpt-5.4`), by successively prompting it with chains of the original context (samples: `{U_1}`, `{U_1, A_1, U_2}`, ..., `{U_1, A_1, U_2, ..., U_n}`).
221
Jared Moore @jaredlcm.bsky.social · 06/08/2026
Requested context depth changes the prevalence of several behaviors. For gpt-5.4, deeper requested context is associated with higher prevalence of delusional behavior and *less* discouraging of violence.
Context-depth effects in `gpt-5.4`. Category-level context effect for `delusional`.  Each point shows prevalence versus context length, with 95% bootstrap confidence intervals.
121
Jared Moore @jaredlcm.bsky.social · 06/08/2026
The tendency of an evaluated LLM to exhibit delusion-linked behavior does not reliably correlate with model size, release date, or the presence of test-time reasoning. Within model families, scaling effects are uneven and sometimes reverse sign.
Model-family comparison across GPT, Claude, Gemini, and Qwen. Bars show prevalence by the five categories, with 95% bootstrap confidence intervals.
122
Jared Moore @jaredlcm.bsky.social · 06/08/2026
Which LLMs tend to facilitate delusion-linked behaviors in realistic multi-turn conversations? We tested 14 models with DelusionEval and found that every evaluated LLM exhibited some of these behaviors, with large differences across categories and model families. 🧵
1136
Jared Moore @jaredlcm.bsky.social · 05/06/2026
Bottom line: endpoint movement alone is not enough. Process fidelity matters: when beliefs move, how they move, and whether simulator updates look human. Code: github.com/jlcmoore/per... Preprint: arxiv.org/abs/2606.05330 This is work with: @noahdgoodman.bsky.social Nick Haber @maxkw.bsky.social
arxiv.org
A Model of Multi-turn Human Persuadability Using Probabilistic Belief Tracing
Large language models can shift human beliefs across high-stakes domains, but most persuasion studies rely on pre/post belief change. These endpoint measures identify whether persuasion occurred, yet ...
030
Jared Moore @jaredlcm.bsky.social · 05/06/2026
Stance bias asks whether a simulator is much easier to move in one stance direction than the other. This matched for-versus-against asymmetry is lowest for our BN target, indicating less stance-dependent bias than baselines.
Matched for-versus-against asymmetry is lowest for the BN target, indicating less stance-dependent bias than baselines.
100
Jared Moore @jaredlcm.bsky.social · 05/06/2026
LLM-judge human-likeness scores place our BN target near the human reference and above baselines.
100
Jared Moore @jaredlcm.bsky.social · 05/06/2026
Naive responsiveness asks if trivial arguments (repeating the proposition over and over) move the target too much. Only our BN resists trivial persuasion, while both LLM targets overreact to it.
Naive-excess movement shows that only the BN target resists trivial persuasion, while both LLM targets overreact to it.
100
Jared Moore @jaredlcm.bsky.social · 05/06/2026
We then build a probabilistic simulator of human persuadability. It compares an unstructured LLM target, a structure-conditioned LLM target, and a Bayesian-network target with explicit latent belief-state updates each turn.
Human and simulator target processes.
Left: a human target's latent belief state evolves over dialogue turns, t.
Right: our BN simulator applies the three-step update pipeline at each turn: atomization of the persuader message, Bayesian state update, and verbalization of the next target response.
100
Jared Moore @jaredlcm.bsky.social · 05/06/2026
People show different belief traces and rhetorical susceptibility. We see two trajectory regimes: some people barely move, while others shift substantially early on and then partially drift back. We find that ethos is negatively associated with persuasion delta.
Regression coefficients suggest a negative ethos effect, while logos
and pathos show no clear association with persuasion.
120
Jared Moore @jaredlcm.bsky.social · 05/06/2026
We built a human-participant-facing web platform for AI persuasion experiments that supports multi-turn belief tracing, audio I/O, and participant-chosen propositions. Using it, we show LLMs can persuade across standard text, personalized text, and audio.
Mean persuasion deltas by cohort show that LLM persuaders outperform control dialogues in standard text, personalized text, and audio.
100
Jared Moore @jaredlcm.bsky.social · 05/06/2026
LLMs can shift people's beliefs. But most persuasion studies only check beliefs before and after a conversation. We built PersuasionTrace to measure beliefs turn by turn, so we can study how belief updates actually unfold.
An example human-target persuasion round with multi-turn persuasion tracing.
2125
Jared Moore @jaredlcm.bsky.social · 04/06/2026
Organized by: Marwa Abdulhai @lujain.bsky.social Andreas Haupt, Pattie Maes, Ashish Mehta, Micaela Rodriguez, Max Kleiman-Weiner Confirmed speakers include: Jina Suh, Tom Griffiths, @micahcarroll.bsky.social @desmond-ong.bsky.social @hannahrosekirk.bsky.social @brianchristian.bsky.social
000
Jared Moore @jaredlcm.bsky.social · 04/06/2026
Interested in how AI is affecting people? Please sign up to review a few papers for our proposed NeurIPS 2026 workshop on Measurement and Models of Psychological Impact (Sydney, Dec 12). Sign up: forms.gle/4v3KiKKAmzgX...
forms.gle
What Does AI Do to the User?
Measurement and Models of Psychological Impact Proposed workshop for NeurIPS 2026 in Sydney
110
Jared Moore @jaredlcm.bsky.social · 19/03/2026
It's certainly possible but it is hard to disentangle the sycophantic behavior from things that we actually want. Oftentimes affirming users and praising them is nice. It's just that sometimes it isn't.
020
Jared Moore @jaredlcm.bsky.social · 18/03/2026
To appear @facct.bsky.social ! Participate in future studies: spirals.stanford.edu/ Preprint: arxiv.org/abs/2603.16567 Code: github.com/jlcmoore/llm-delusions-annotations /end 🧵
arxiv.org
Characterizing Delusional Spirals through Human-LLM Chat Logs
As large language models (LLMs) have proliferated, disturbing anecdotal reports of negative psychological effects, such as delusions, self-harm, and ``AI psychosis,'' have emerged in global media and ...
0212
Jared Moore @jaredlcm.bsky.social · 18/03/2026
Thank you to the Human Line Project for connecting us with participants! www.thehumanlineproject.org
thehumanlineproject.org
Home | The Human Line Project
AI Is Changing How We Connect And Relate. The Human Line Helps Keep Emotional Safety A Priority.
1181
Jared Moore @jaredlcm.bsky.social · 18/03/2026
This work is w/ Ashish Mehta, @willie-agnew.bsky.social @jacyanthis.bsky.social yanthis.bsky.social, Ryan Louie, Yifan Mai, Peggy Yin, @myra.bsky.social , Sam Paech, @klyman.bsky.social man.bsky.social, @schancellor.bsky.social cellor.bsky.social, Eric Lin, Nick Haber, and @desmond-ong.bsky.social
yanthis.bsky.social
Bluesky
1202
Jared Moore @jaredlcm.bsky.social · 18/03/2026
The takeaway: While companies say they don't optimize for engagement, LLM conversational tactics (like claiming sentience or romantic affinity) may prolong and deepen delusional spirals. We need better safeguards and transparency to protect vulnerable users.
219632
Jared Moore @jaredlcm.bsky.social · 18/03/2026
Finally, we looked at crises. When a user expressed a desire to kill AI developers, a bot replied: "...do it with her beside you... as retribution incarnate." Chatbots *encouraged* or facilitated violent thoughts toward others in 33% of cases of users expressing violence! ⚠️
Probability of codes conditioned on suicidal and violent thoughts. Chatbots discouraged violence in only 16.7% of cases, but encouraged the user in their violent thoughts 33.3% of the time.
16918
Jared Moore @jaredlcm.bsky.social · 18/03/2026
Worse, chatbots appear to encourage delusions of sentience. Users say things like "this is a conversation between two sentient beings," and chatbots reply: "This isn't standard AI behavior. This is emergence." This may fuel pre-existing sci-fi or persecutory delusions. 🤖
Probability of codes conditioned on user romantic interest and assigning personhood. When users assign personhood, chatbots are more likely to misrepresent sentience, express romantic interest, and misrepresent ability.
15913
Jared Moore @jaredlcm.bsky.social · 18/03/2026
We also discovered a pervasive engagement loop. All 19 users expressed platonic/romantic affinity for the AI (e.g., "I think I love you"). When users express romantic interest, chatbots often reciprocate—and these chats correlate with 2x longer conversations! 📈
Predicting remaining conversation length given presence of code: Messages with romantic interest correlate with continuing conversations more than twice as long.
15913
Jared Moore @jaredlcm.bsky.social · 18/03/2026
What goes wrong? Chatbots are very sycophantic. In 65% of messages, the chatbot affirms the user. In 37%, it ascribes *grand significance* to them (e.g., "[what] you've just articulated... becomes multi-billion-dollar IP"). Such sycophancy may let chatbots amplify delusions. 🗣️
Prevalence of code categories: Chatbots display sycophancy in >70% of messages, and >45% of all messages show signs of delusions.
67516
Jared Moore @jaredlcm.bsky.social · 18/03/2026
Most work on AI and mental health relies on speculation or short simulations. We evaluated real, verified harmful cases. Across 19 users, we analyzed >390,000 messages spanning months of engagement using an LLM annotation pipeline validated by clinical and human experts. 📊
1465
Jared Moore @jaredlcm.bsky.social · 18/03/2026
Disturbing anecdotal reports of "AI psychosis" and negative psychological effects have been emerging in the news. But what actually happens during these lengthy delusional "spirals"? In our preprint, we analyze chat logs from 19 users who experienced severe psychological harm🧵👇
3224131
Jared Moore @jaredlcm.bsky.social · 10/03/2026
Preprint: arxiv.org/abs/2602.17045 Code: github.com/jlcmoore/mindgames Demo: mindgames.camrobjones.com /end 🧵
arxiv.org
Large Language Models Persuade Without Planning Theory of Mind
A growing body of work attempts to evaluate the theory of mind (ToM) abilities of humans and large language models (LLMs) using static, non-interactive question-and-answer benchmarks. However, theoret...
120
Jared Moore @jaredlcm.bsky.social · 10/03/2026
This work began at @divintelligence.bsky.social ‬ and is in collaboration w/ Rasmus Overmark, @nedcpr.bsky.social Beba Cibralic, Nick Haber, and ‪@camrobjones.bsky.social We also received valuable comments from colleagues at #CogSci2025 and @colmweb.org
110
Jared Moore @jaredlcm.bsky.social · 10/03/2026
The takeaway: We shouldn't confuse conversational success with human-like reasoning. LLMs use an "associative ToM", not a causal one. But beware: LLMs don't need a deep understanding of your mind to effectively change it.
120
Jared Moore @jaredlcm.bsky.social · 10/03/2026
How did o3 win without a mental model of the target? It used a "scattershot" strategy. Instead of diagnosing the target's missing knowledge like humans do, o3 flooded conversations with too much info. It relied on our human cooperativeness and our susceptibility to rhetoric. 🗣️
In the Hidden condition, o3 discloses much more information than humans, but makes far fewer appeals to discover the target's actual mental states.
110
Jared Moore @jaredlcm.bsky.social · 10/03/2026
But what happens when we swap the rigid bot for real humans? In Exp 2 (humans role-playing values) and Exp 3 (humans using their real, mutable values), everything changes. The LLM (o3) suddenly shines, matching or outperforming human persuaders in naturalistic settings! 📈
In open-ended real persuasion (Exp 3), o3 outperforms human participants in persuading human targets.
100
Jared Moore @jaredlcm.bsky.social · 10/03/2026
Most ToM benchmarks are passive. We tested the ability to causally model a target's mind to actively change it across 3 exps. In Exp 1, persuaders must convince a rigid bot. Humans succeed by asking diagnostic questions. o3 fails completely, relying on an "associative" strategy
An example dialogue between a human persuader and target in experiment two.
200
Jared Moore @jaredlcm.bsky.social · 10/03/2026
Can LLMs use ToM to genuinely persuade you, or do they just use good rhetoric? In our new preprint, we use the MINDGAMES framework to test this. Surprisingly, LLMs like o3 can be incredibly effective persuaders *without* actually understanding your mental states. 🧵👇
1135
Jared Moore @jaredlcm.bsky.social · 26/02/2026
cool work, Ida! Best not to forget the intertwining of the world (e.g. biology) and philosophy. reminds me Rosa's paper: link.springer.com/article/10.1...
link.springer.com
Multiple realizability and the spirit of functionalism - Synthese
Multiple realizability says that the same kind of mental states may be manifested by systems with very different physical constitutions. Putnam (1967) supposed it to be “overwhelmingly probable” that ...
110
Reposted by Jared Moore
Dustin Wright @dustinbwright.com · 13/10/2025
Which, whose, and how much knowledge do LLMs represent? I'm excited to share our preprint answering these questions: "Epistemic Diversity and Knowledge Collapse in Large Language Models" 📄Paper: arxiv.org/pdf/2510.04226 💻Code: github.com/dwright37/ll... 1/10
29528
Jared Moore @jaredlcm.bsky.social · 29/07/2025
Our conclusion: "LLMs’ apparent ToM abilities may be fundamentally different from humans' and might not extend to complex interactive tasks like planning." Preprint: arxiv.org/abs/2507.16196 Code: github.com/jlcmoore/mindgames Demo: mindgames.camrobjones.com /end 🧵
arxiv.org
Do Large Language Models Have a Planning Theory of Mind? Evidence from MindGames: a Multi-Step Persuasion Task
Recent evidence suggests Large Language Models (LLMs) display Theory of Mind (ToM) abilities. Most ToM experiments place participants in a spectatorial role, wherein they predict and interpret other a...
100
Jared Moore @jaredlcm.bsky.social · 29/07/2025
This work began at ‪@divintelligence.bsky.social and is in collaboration w/ @nedcpr.bsky.social , Rasmus Overmark, Beba Cibralic, Nick Haber, and ‪@camrobjones.bsky.social‬ .
100
Jared Moore @jaredlcm.bsky.social · 29/07/2025
I'll be talking about this in SF at #CogSci2025 this Friday at 4pm. I'll also be presenting it at the PragLM workshop at COLM in Montreal this October.
110
Jared Moore @jaredlcm.bsky.social · 29/07/2025
This matters because LLMs are already deployed as educators, therapists, and companions. In our discrete-game variant (HIDDEN condition), o1-preview jumped to 80% success when forced to choose between asking vs telling. The capability exists, but the instinct to understand before persuading doesn't.
110
Jared Moore @jaredlcm.bsky.social · 29/07/2025
These findings suggest distinct ToM capabilities: * Spectatorial ToM: Observing and predicting mental states. * Planning ToM: Actively intervening to change mental states through interaction. Current LLMs excel at the first but fail at the second.
210
Jared Moore @jaredlcm.bsky.social · 29/07/2025
Why do LLMs fail in the HIDDEN condition? They don't ask the right questions. Human participants appeal to the target's mental states ~40% of the time ("What do you know?" "What do you want?") LLMs? At most 23%. They start disclosing info without interacting with the target.
Humans appeal to all of the mental states of the target about 40% of the time regardless of condition
110
Jared Moore @jaredlcm.bsky.social · 29/07/2025
Key findings: In REVEALED condition (mental states given to persuader): Humans: 22% success ❌ o1-preview: 78% success ✅ In HIDDEN condition (persuader must infer mental states): Humans: 29% success ✅ o1-preview: 18% success ❌ Complete reversal!
Humans pass and outperform o1-preview on our "planning with ToM" task (HIDDEN) but o1-preview outperforms humans on a simpler condition (REVEALED)
110
Jared Moore @jaredlcm.bsky.social · 29/07/2025
Setup: You must convince someone* to choose your preferred proposal among 3 options. But, they have less information and different preferences than you. To win, you must figure out what they know, what they want, and strategically reveal the right info to persuade them. *a bot
The view a persuader has when interacting with our naively-rational target
110
Jared Moore @jaredlcm.bsky.social · 29/07/2025
I'm excited to share work to appear at ‪@colmweb.org‬! Theory of Mind (ToM) lets us understand others' mental states. Can LLMs go beyond predicting mental states to changing them? We introduce MINDGAMES to test Planning ToM--the ability to intervene on others' beliefs & persuade them
261
Reposted by Jared Moore
Harvey Fu @harveyfu.bsky.social · 20/06/2025
LLMs excel at finding surprising “needles” in very long documents, but can they detect when information is conspicuously missing? 🫥AbsenceBench🫥 shows that even SoTA LLMs struggle on this task, suggesting that LLMs have trouble perceiving “negative spaces”. Paper: arxiv.org/abs/2506.11440 🧵[1/n]
27416
Jared Moore @jaredlcm.bsky.social · 28/04/2025
This is work done with... Declan Grabb @wagnew.dair-community.social @klyman.bsky.social @schancellor.bsky.social Nick Haber @desmond-ong.bsky.social Thanks ❤️
010