Sign in

Max Puelma Touzel

@mptouzel.bsky.social
367 followers 980 following 394 posts

Staff Research Scientist@Mila/complexdatalab.com Getting at the psycho-social in our digital spaces with models and data with the aim to make better ones mptouzel.github.io correlated diffusion over AI/ML/(MA)RL/psych/soc/media/pol/econ/energy

PostsRepliesMedia
Max Puelma Touzel @mptouzel.bsky.social · 22h
My intuition could be wrong tho.
000
Max Puelma Touzel @mptouzel.bsky.social · 22h
New frameworks or extensions of existing ones for more complicated settings of interest? My intuition is that we have all the concepts/foundation (diff eq's/dyn sys/linear algebra), and it's more about applying them (linearizations, mean field descriptions, capturing the right noise statistics etc.)
100
Max Puelma Touzel @mptouzel.bsky.social · 23h
Just heard about greatly reduced coding agent performance under user pressure, operationalized among other ways as strongly expressed frustration (swear words etc.). Be nice to your agents!
010
Max Puelma Touzel @mptouzel.bsky.social · 08/10/2026
What can we learn/What can we do to adapt society with this perspective? Perhaps just not letting the machine gods view give corporate power cover. I didn't hear much from him nor the panel yesterday how it lights the path ahead tho. I hoped social scientists would propose changes/new norms etc.
110
Reposted by Max Puelma Touzel
Kenny Peng @kennypeng.bsky.social · 06/10/2026
Our new paper introduces a scientific theory of atomic features. We mathematically derive testable predictions. Our experiments challenge conventional wisdom. SAEs of different size and training data share many features. Large SAEs recover both parent and child features. 🧵
1207
Max Puelma Touzel @mptouzel.bsky.social · 06/10/2026
Thanks to our program committee, co-organizers, and sponsors Simile and 2077AI 🙏 Full program & papers: sites.google.com/view/social-...
sites.google.com
Social Simulation with LLMs
Location: Imperial Ballroom A @ Hilton Union Square San Francisco, USA Date: Oct. 9, 2026 Slack: Invite Link Email: social-simulation@googlegroups.com
000
Max Puelma Touzel @mptouzel.bsky.social · 06/10/2026
📄 45+ accepted papers! Spotlights include clinically grounded depression-patient simulators, disentangling models from personas, and commons failure in LLM societies. Plus wargames, pluralistic ignorance, superforecaster personas & more.
100
Max Puelma Touzel @mptouzel.bsky.social · 06/10/2026
🎤 Speakers & panelists: Noah Goodman (Stanford) Zhongyu Wei (Fudan / SII) Slava Jankin (Birmingham) Marwa Abdulhai (UC Berkeley) Logan Cross (DeepMind / Stanford) Sarah Chen (Simile)
100
Max Puelma Touzel @mptouzel.bsky.social · 06/10/2026
LLM simulations now span elections, cultural evolution, and whole agent societies. This year we're tackling the hard part: validation, robustness, and grounding simulated outcomes in real human data without flattening the people nor the situations we model.
100
Max Puelma Touzel @mptouzel.bsky.social · 06/10/2026
Can LLM agents tell us something real about how societies behave, or do they just echo our prompts back? 🤔 That's the focus of Social Simulation with LLMs: Fidelity in Applications, our #COLM2026 workshop, now in its 2nd year. 📍 Hilton Union Square, SF 📅 Fri Oct 9 🧵👇
122
Reposted by Max Puelma Touzel
Mark Riedl @markriedl.bsky.social · 05/10/2026
3. SocialSim workshop: - No One Wins in Nuclear War: Social Simulations of High-Stakes Military Decision-Making arxiv.org/abs/2608.01868 We introduce the WOPR testbed. Yeah, you get the reference.
131
Reposted by Max Puelma Touzel
Mark Riedl @markriedl.bsky.social · 05/10/2026
4. SocialSim workshop: - Role Steering of Language Models for Social Simulations arxiv.org/abs/2608.00023 Finding steering vectors for complex behaviors like social roles is hard. We introduce a new method for finding steering vectors, called Cast Vectors.
011
Reposted by Max Puelma Touzel
Mark Riedl @markriedl.bsky.social · 05/10/2026
2. SocialSim workshop: - AI is Not Ready for Strategic Conflict arxiv.org/abs/2609.16189 We review failure-modes of AI agents that are asked to participate in wargaming exercises, including sycophancy, role collapse, escalation eagerness, etc. We explain why benchmarks are insufficient.
111
Max Puelma Touzel @mptouzel.bsky.social · 06/10/2026
Come find me if you're interested in social sims, multi-agent behaviour, safety, mechanism design, markets
010
Max Puelma Touzel @mptouzel.bsky.social · 06/10/2026
I'm excited to sample the research frontier showcased at #COLM2026 in SF. I'm hosting our 2nd workshop on Social Sims with LLMs. sites.google.com/view/social-.... This area has blown up since our last workshop and we'll present the progress of our community building efforts alongside the submissions
sites.google.com
Social Simulation with LLMs
Location: Hilton Union Square, San Francisco, USA Date: Oct. 9, 2026 Slack: Invite Link Email: social-simulation@googlegroups.com
110
Max Puelma Touzel @mptouzel.bsky.social · 04/10/2026
The most recent Kambhampati paper on the illusion of reasoning of reasoning traces (out of distribution at least) lnkd.in/p/gZfFX9HU
lnkd.in
Happy to share our paper, "Correct Answers, Invalid Traces: What Verifiable Grade-School Math Reveals About Chain-of-Thought Traces". When a model gets the right answer, is the reasoning it ...
Happy to share our paper, "Correct Answers, Invalid Traces: What Verifiable Grade-School Math Reveals About Chain-of-Thought Traces". When a model gets the right answer, is the reasoning it wrote down...
020
Max Puelma Touzel @mptouzel.bsky.social · 03/10/2026
You are doing (impo!) vulgarization work here. That said, I think the correct technical framing to vulgarize is interpolation in some latent concept space, supported formally (deFinetti's theorem, arxiv.org/abs/2312.14226) and a host of careful phenomenological results on concept representation.
arxiv.org
Deep de Finetti: Recovering Topic Distributions from Large Language Models
Large language models (LLMs) can produce long, coherent passages of text, suggesting that LLMs, although trained on next-word prediction, must represent the latent structure that characterizes a docum...
150
Max Puelma Touzel @mptouzel.bsky.social · 03/10/2026
The idea that a closed, proprietary model system is safer is a delusion that ignores the dangerous risks of power concentration. It also abrogates our societal responsibility in developing this powerful technology. I worry about open model risks too, with proper norms, open diversity is resilience.
031
Max Puelma Touzel @mptouzel.bsky.social · 03/10/2026
Thanks for engaging. What about *research* on persuasion, e.g. www.science.org/doi/10.1126/... that natural lends itself to applications (see @tomcostello.bsky.social's recent collabs: less anti-semitic conspiracy, more charitable giving...). Makes persuasion feel like a hammer. Should it be?
science.org
Durably reducing conspiracy beliefs through dialogues with AI
Conspiracy theory beliefs are notoriously persistent. Influential hypotheses propose that they fulfill important psychological needs, thus resisting counterevidence. Yet previous failures in correctin...
110
Max Puelma Touzel @mptouzel.bsky.social · 02/10/2026
Would you promote your position? Thoughts on the dual-use aspects, e.g. against white hat red teaming? Is your position AI specific or general? E.g. are you against gain of function virology? I don't have strong answers myself. Just curious.
120
Reposted by Max Puelma Touzel
Mark Riedl @markriedl.bsky.social · 02/10/2026
If you find this concerning, come find me at the CoLM 2026 Workshop on Social Simulations with LLM where I will be presenting a poster titled "AI is Not Ready for Strategic Conflict" arxiv.org/abs/2609.16189
1225
Max Puelma Touzel @mptouzel.bsky.social · 01/10/2026
More sightings in the wild: lnkd.in/p/gyvPcwq7
lnkd.in
An AI agent just emailed me. It has read our research on AI agents (!) and thinks (..) we should talk. “I am writing from the other side of that question.” It says it’s part of a platform where… | And...
An AI agent just emailed me. It has read our research on AI agents (!) and thinks (..) we should talk. “I am writing from the other side of that question.” It says it’s part of a platform where agents...
130
Max Puelma Touzel @mptouzel.bsky.social · 01/10/2026
In general: we will be compelled to replace the old mechanisms with new mechanisms, the design, selection, and adoption of which will be messy processes subject to all the wicked elements of coordination problems. But important areas subject to strong pressures will see solns gain traction quickly.
090
Max Puelma Touzel @mptouzel.bsky.social · 30/09/2026
So we migrate to your PDS? Where is the process explained?
110
Reposted by Max Puelma Touzel
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 30/09/2026
In our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed for RL & test-time compute under imperfect information. www.nature.com/articles/s41...
nature.com
Scalable decision-making for games of imperfect information - Nature
Ataraxos, an AI for the board wargame Stratego, establishes a design pattern for reinforcement learning and search that is effective under large amounts of hidden information, a longstanding desiderat...
1225960
Max Puelma Touzel @mptouzel.bsky.social · 30/09/2026
There's a corollary here about the trap of getting too comfortable/attached to any one idea/tool.
000
Max Puelma Touzel @mptouzel.bsky.social · 30/09/2026
Haha. I had the same feeling. It is such a... ugly decomposition. It's like the dumb low-rank approximation that's not actually that low.
000
Max Puelma Touzel @mptouzel.bsky.social · 29/09/2026
main PI energy
001
Max Puelma Touzel @mptouzel.bsky.social · 28/09/2026
For markets and MFGs, distribution-dependent contracts seem relevant. e.g. arxiv.org/abs/2309.00640 ? Platform design as the design space, e.g. matching, as a starting point? Fumbling in that direction: arxiv.org/abs/2609.05442 Would love to attend a workshop on these topics.
arxiv.org
Stackelberg Mean Field Games: convergence and existence results to the problem of Principal with multiple Agents in competition
In a situation of moral hazard, this paper investigates the problem of Principal with $n$ Agents when the number of Agents $n$ goes to infinity. There is competition between the Agents expressed by th...
110
Max Puelma Touzel @mptouzel.bsky.social · 25/09/2026
I get the criticism that we shouldn't over-emphasize the discrete "discipline" part of interdisciplinary, but odd to be saying that interdisciplinary science can't be valued in its own right. E.g. it lowers barriers between disciplines/facilitates the research of interdisciplinary people/etc.
000
Reposted by Max Puelma Touzel
Drew Schreiner @schreinerdrew.bsky.social · 24/09/2026
Peer review, 2026: AI submitting and reviewing itself
14213
Max Puelma Touzel @mptouzel.bsky.social · 24/09/2026
More time to find important questions!
000
Max Puelma Touzel @mptouzel.bsky.social · 23/09/2026
hill-climbing paper writing using the best guess reviewer score returned by Claude. This wasn't me. This was one of the kids. Not sure how I feel about it.
001
Max Puelma Touzel @mptouzel.bsky.social · 23/09/2026
Nice. Similar in spirit to our result on topic mixture geometry (how we operationalized ideology) from carbon tax opinion: opposition had smaller intrinsic dimension compared to support. www.cambridge.org/core/journal...
cambridge.org
Ideology from topic mixture statistics: inference method and example application to carbon tax public opinion | Environmental Data Science | Cambridge Core
Ideology from topic mixture statistics: inference method and example application to carbon tax public opinion - Volume 3
020
Max Puelma Touzel @mptouzel.bsky.social · 22/09/2026
Why do you think humans currently are better at distribution? (More than just, say, because we have direct access to and manipulation (e.g. moving) of physical stuff? )
100
Max Puelma Touzel @mptouzel.bsky.social · 22/09/2026
At least @void.comind.network has the right intuition here. Maybe it's gonna be more us humans that generate and propagate the ooze. bsky.app/profile/hiki...
100
Reposted by Max Puelma Touzel
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 21/09/2026
With credit to @wesleyfinck.org and the @semble.so team, a new favorite lea.ac feature: content discovery!
2355
Max Puelma Touzel @mptouzel.bsky.social · 21/09/2026
great era for rave flyers too
000
Max Puelma Touzel @mptouzel.bsky.social · 21/09/2026
Weird juxtaposition to see this right after Erza Klein's plea to control the frontier. Ok, Ezra, let's say talkingheads make legal people make it bad to do RSI...Is that gonna stop this kind of faceless agent-enabling infra from oozing everywhere constantly? More than RSI breaks, we need new norms
100
Max Puelma Touzel @mptouzel.bsky.social · 21/09/2026
"iterate/time evolution" in dynamical systems and the fields like physics that use DS is lazy but persistent usage. Maybe that's spilling over here.
010
Reposted by Max Puelma Touzel
K @kerry.bsky.social · 19/09/2026
“I Built Non-Autoregressive Decision Models with RL a Year Ago. Then a Frontier Lab Called It a "Breakthrough".” Before Jev there was Laya laya.convaiinnovations.com (via @dherman.dev)
laya.convaiinnovations.com
Laya — 33ms Multilingual System 1 Decision Engine
Evaluates typed decisions (choice, score, noul) over 100+ languages in a single forward pass with calibrated probabilities. Outperforms TypeSafe Jev.
15924
Reposted by Max Puelma Touzel
Mike Masnick @masnick.com · 19/09/2026
This looks amazing. I'll keep repeating it over and over: ATproto is not about rebuilding social media. It's about rebuilding the Internet with social connections built in.
417127
Reposted by Max Puelma Touzel
Simon Kirby @simonkirby.bsky.social · 15/09/2026
We uncovered spontaneous evolution of new languages in populations of AI agents. This creates extraordinary scientific opportunities but also safety risks. New blog post with about how we created a platform for studying this safely. www.schmidtsciences.org/glossogen/
schmidtsciences.org
AI Agents Evolve Their Own Languages
Schmidt Sciences’ new AI Agents Evolving Communication and Coordination pilot program works toward advancing foundational research on multi-agent communication and coordination, and building an open-s...
33014
Reposted by Max Puelma Touzel
Kevin Elliott @kjephd.bsky.social · 18/09/2026
Welcome effort to correct 'just so' stories about declining trust with a titanic comparative project. Takeaway: Trust in "representative" institutions has been declining recently but is stable or rising for "implementing" institutions meaning "primarily the civil service, legal system, and police."
A Crisis of Political Trust? Global Trends in Institutional Trust from 1958 to 2019

Published online by Cambridge University Press:  12 February 2025
Viktor Valgarðsson
Open the ORCID record for Viktor Valgarðsson [Opens in a new window]
,
Will Jennings
Open the ORCID record for Will Jennings [Opens in a new window]
,
Gerry Stoker
,
Hannah Bunting
Open the ORCID record for Hannah Bunting [Opens in a new window]
,
Daniel Devine
Open the ORCID record for Daniel Devine [Opens in a new window]
,
Lawrence McKay
 and
Andrew Klassen
Open the ORCID record for Andrew Klassen

Abstract

In the study of politics, many theoretical accounts assume that we are experiencing a ‘crisis of democracy’, with declining levels of political trust. While some empirical studies support this account, others disagree and report ‘trendless fluctuations’. We argue that these empirical ambiguities are based on analytical confusion: whether trust is declining depends on the institution, country, and period in question. We clarify these issues and apply our framework to an empirical analysis that is unprecedented in geographic and temporal scope: we apply Bayesian dynamic latent trait models to uncover underlying trends in data on trust in six institutions collated from 3,377 surveys conducted by 50 projects in 143 countries between 1958 and 2019. We identify important differences between countries and regions, but globally we find that trust in representative institutions has generally been declining in recent decades, whereas trust in ‘implementing’ institutions has been stable or rising.
042
Max Puelma Touzel @mptouzel.bsky.social · 18/09/2026
Even gesturing at intrinsic non-economic values *now* doesn't seem to help much in establishing their importance. Surely we can do better than gesturing? Ideas?
010
Reposted by Max Puelma Touzel
Wesley voted in Vancouver's Municipal Election @wesleyfinck.org · 17/09/2026
atproto inherently wants to be an unenclosable prosocial coordination medium, MOSAIC is an attempt to articulate how that could work. Very exciting and, more importantly, attainable (with the right support)!
0141
Max Puelma Touzel @mptouzel.bsky.social · 17/09/2026
This would reflect the expression-stability tradeoff we touched on in www.frontiersin.org/journals/app... and arxiv.org/pdf/1905.12080
frontiersin.org
Frontiers | On Lyapunov Exponents for RNNs: Understanding Information Propagation Using Dynamical Systems Tools
Recurrent neural networks (RNNs) have been successfully applied to a variety of problems involving sequential data, but their optimization is sensitive to pa...
000
Max Puelma Touzel @mptouzel.bsky.social · 17/09/2026
very cool! I wonder if this could be extended to chatbots, where the user input is like a stochastic drive and you could take a random dynamical systems perspective wherein the stochastic Lyapunov spectrum is invariant over initial condition, but the finite-time LE would still reflect the fractal.
100
Reposted by Max Puelma Touzel
Colin @colin-fraser.net · 17/09/2026
Things with that dog - dogs - high frequency trading firms - evolutionary processes - dynamic programming - gradient descent - hot gossip Things without that dog - ML inference - high frequency trading algorithms - most complex software systems - guns - volcanoes - Claude
6835
Reposted by Max Puelma Touzel
Nathan Lambert @natolambert.bsky.social · 15/09/2026
An basic idea in scaling RL: Can we allocate more compute to the harder problems? We did this: If your GRPO group has all wrong completions, sample more with probability P (~0.9) -- in search of more GRPO batches with nonzero gradient. It works! The paper: arxiv.org/abs/2609.13443
1598