Sign in

benaslater.bsky.social

@benaslater.bsky.social
1 followers 1 following 7 posts

PhD Student @ University of Cambridge, interested in AI evaluation

PostsRepliesMedia
benaslater.bsky.social @benaslater.bsky.social · 25/09/2026
This work has been accepted to NeurIPS 2026! If you're interested in chatting about evaluating AI theory of mind, planning, safety, or anything else come and find me at NeurIPS in December
000
benaslater.bsky.social @benaslater.bsky.social · 10/07/2026
Thank you to coauthors Matteo G. Mecattaf, @lucycheke.bsky.social , John Burden and Winnie Street. This work was completed partially in the ERA AI safety fellowship, and partially during my PhD: Take a look at the paper at arxiv.org/abs/2606.31916!
arxiv.org
Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action
Theory of Mind (ToM) benchmarks for Large Language Models (LLMs) typically rely on passive question-answering formats, but the deployment of LLMs in increasingly agentic and autonomous forms demands n...
111
benaslater.bsky.social @benaslater.bsky.social · 10/07/2026
There's a lot further to go to understand the full extent of NCP-ToM: Real scenarios will have much more complexity, both in the environment and in those being persuaded. We hope our work inspires others evaluating social intelligence in models to consider the NCP-ToM aspects of their scenarios.
121
benaslater.bsky.social @benaslater.bsky.social · 10/07/2026
Our design also allowed us to pair each NCP-ToM task (Agentic) with a conventional ToM task (Q&A). We thought that the Agentic tasks might be strictly more difficult than the Q&A ones, but that broke down for GPT-5, which had lots of cases where it failed the Q&A task but passed the Agentic.
111
benaslater.bsky.social @benaslater.bsky.social · 10/07/2026
We evaluated six frontier models, including GPT-5, Gemini 2.5 Pro, Claude models, and a cohort of human participants, across 600 task instances. GPT-5 was the only model to outperform human participants on our task. However, humans were still more robust to changes in the context than models.
111
benaslater.bsky.social @benaslater.bsky.social · 10/07/2026
We inverted conventional ToM tasks to task the model with inducing belief states in simulated characters, by curating what they observe. Achieving many beliefs simultaneously requires a lot of planning, so we call the capability we're testing Non-Conversational Planning Theory of Mind (NCP-ToM).
111
benaslater.bsky.social @benaslater.bsky.social · 10/07/2026
New work on AI Theory of Mind: We ask: *How capable are models at inducing belief states without using conversation?* This is important to track as it unlocks many AI agent use-cases, but also potential harms. In our task, recent generations of models exceeded human performance 🧵
121