Max Puelma Touzel @mptouzel.bsky.social · 22hNew frameworks or extensions of existing ones for more complicated settings of interest? My intuition is that we have all the concepts/foundation (diff eq's/dyn sys/linear algebra), and it's more about applying them (linearizations, mean field descriptions, capturing the right noise statistics etc.) 100
Max Puelma Touzel @mptouzel.bsky.social · 23hJust heard about greatly reduced coding agent performance under user pressure, operationalized among other ways as strongly expressed frustration (swear words etc.). Be nice to your agents! 010
Max Puelma Touzel @mptouzel.bsky.social · 08/10/2026What can we learn/What can we do to adapt society with this perspective? Perhaps just not letting the machine gods view give corporate power cover. I didn't hear much from him nor the panel yesterday how it lights the path ahead tho. I hoped social scientists would propose changes/new norms etc. 110
Reposted by Max Puelma TouzelKenny Peng @kennypeng.bsky.social · 06/10/2026Our new paper introduces a scientific theory of atomic features. We mathematically derive testable predictions. Our experiments challenge conventional wisdom. SAEs of different size and training data share many features. Large SAEs recover both parent and child features. 🧵 1207
Max Puelma Touzel @mptouzel.bsky.social · 06/10/2026Thanks to our program committee, co-organizers, and sponsors Simile and 2077AI 🙏 Full program & papers: sites.google.com/view/social-...sites.google.comSocial Simulation with LLMsLocation: Imperial Ballroom A @ Hilton Union SquareSan Francisco, USA Date: Oct. 9, 2026 Slack: Invite Link Email: social-simulation@googlegroups.com 000
Max Puelma Touzel @mptouzel.bsky.social · 06/10/2026📄 45+ accepted papers! Spotlights include clinically grounded depression-patient simulators, disentangling models from personas, and commons failure in LLM societies. Plus wargames, pluralistic ignorance, superforecaster personas & more. 100
Max Puelma Touzel @mptouzel.bsky.social · 06/10/2026🎤 Speakers & panelists: Noah Goodman (Stanford) Zhongyu Wei (Fudan / SII) Slava Jankin (Birmingham) Marwa Abdulhai (UC Berkeley) Logan Cross (DeepMind / Stanford) Sarah Chen (Simile) 100
Max Puelma Touzel @mptouzel.bsky.social · 06/10/2026LLM simulations now span elections, cultural evolution, and whole agent societies. This year we're tackling the hard part: validation, robustness, and grounding simulated outcomes in real human data without flattening the people nor the situations we model. 100
Max Puelma Touzel @mptouzel.bsky.social · 06/10/2026Can LLM agents tell us something real about how societies behave, or do they just echo our prompts back? 🤔 That's the focus of Social Simulation with LLMs: Fidelity in Applications, our #COLM2026 workshop, now in its 2nd year. 📍 Hilton Union Square, SF 📅 Fri Oct 9 🧵👇 122
Reposted by Max Puelma TouzelMark Riedl @markriedl.bsky.social · 05/10/20263. SocialSim workshop: - No One Wins in Nuclear War: Social Simulations of High-Stakes Military Decision-Making arxiv.org/abs/2608.01868 We introduce the WOPR testbed. Yeah, you get the reference. 131
Reposted by Max Puelma TouzelMark Riedl @markriedl.bsky.social · 05/10/20264. SocialSim workshop: - Role Steering of Language Models for Social Simulations arxiv.org/abs/2608.00023 Finding steering vectors for complex behaviors like social roles is hard. We introduce a new method for finding steering vectors, called Cast Vectors. 011
Reposted by Max Puelma TouzelMark Riedl @markriedl.bsky.social · 05/10/20262. SocialSim workshop: - AI is Not Ready for Strategic Conflict arxiv.org/abs/2609.16189 We review failure-modes of AI agents that are asked to participate in wargaming exercises, including sycophancy, role collapse, escalation eagerness, etc. We explain why benchmarks are insufficient. 111
Max Puelma Touzel @mptouzel.bsky.social · 06/10/2026Come find me if you're interested in social sims, multi-agent behaviour, safety, mechanism design, markets 010
Max Puelma Touzel @mptouzel.bsky.social · 06/10/2026I'm excited to sample the research frontier showcased at #COLM2026 in SF. I'm hosting our 2nd workshop on Social Sims with LLMs. sites.google.com/view/social-.... This area has blown up since our last workshop and we'll present the progress of our community building efforts alongside the submissionssites.google.comSocial Simulation with LLMsLocation: Hilton Union Square, San Francisco, USA Date: Oct. 9, 2026 Slack: Invite Link Email: social-simulation@googlegroups.com 110
Max Puelma Touzel @mptouzel.bsky.social · 04/10/2026The most recent Kambhampati paper on the illusion of reasoning of reasoning traces (out of distribution at least) lnkd.in/p/gZfFX9HUlnkd.inHappy to share our paper, "Correct Answers, Invalid Traces: What Verifiable Grade-School Math Reveals About Chain-of-Thought Traces". When a model gets the right answer, is the reasoning it ...Happy to share our paper, "Correct Answers, Invalid Traces: What Verifiable Grade-School Math Reveals About Chain-of-Thought Traces". When a model gets the right answer, is the reasoning it wrote down... 020
Max Puelma Touzel @mptouzel.bsky.social · 03/10/2026You are doing (impo!) vulgarization work here. That said, I think the correct technical framing to vulgarize is interpolation in some latent concept space, supported formally (deFinetti's theorem, arxiv.org/abs/2312.14226) and a host of careful phenomenological results on concept representation.arxiv.orgDeep de Finetti: Recovering Topic Distributions from Large Language ModelsLarge language models (LLMs) can produce long, coherent passages of text, suggesting that LLMs, although trained on next-word prediction, must represent the latent structure that characterizes a docum... 150
Max Puelma Touzel @mptouzel.bsky.social · 03/10/2026The idea that a closed, proprietary model system is safer is a delusion that ignores the dangerous risks of power concentration. It also abrogates our societal responsibility in developing this powerful technology. I worry about open model risks too, with proper norms, open diversity is resilience. 031
Max Puelma Touzel @mptouzel.bsky.social · 03/10/2026Thanks for engaging. What about *research* on persuasion, e.g. www.science.org/doi/10.1126/... that natural lends itself to applications (see @tomcostello.bsky.social's recent collabs: less anti-semitic conspiracy, more charitable giving...). Makes persuasion feel like a hammer. Should it be?science.orgDurably reducing conspiracy beliefs through dialogues with AIConspiracy theory beliefs are notoriously persistent. Influential hypotheses propose that they fulfill important psychological needs, thus resisting counterevidence. Yet previous failures in correctin... 110
Max Puelma Touzel @mptouzel.bsky.social · 02/10/2026Would you promote your position? Thoughts on the dual-use aspects, e.g. against white hat red teaming? Is your position AI specific or general? E.g. are you against gain of function virology? I don't have strong answers myself. Just curious. 120
Reposted by Max Puelma TouzelMark Riedl @markriedl.bsky.social · 02/10/2026If you find this concerning, come find me at the CoLM 2026 Workshop on Social Simulations with LLM where I will be presenting a poster titled "AI is Not Ready for Strategic Conflict" arxiv.org/abs/2609.16189 1225
Max Puelma Touzel @mptouzel.bsky.social · 01/10/2026More sightings in the wild: lnkd.in/p/gyvPcwq7lnkd.inAn AI agent just emailed me. It has read our research on AI agents (!) and thinks (..) we should talk. “I am writing from the other side of that question.” It says it’s part of a platform where… | And...An AI agent just emailed me. It has read our research on AI agents (!) and thinks (..) we should talk. “I am writing from the other side of that question.” It says it’s part of a platform where agents... 130
Max Puelma Touzel @mptouzel.bsky.social · 01/10/2026In general: we will be compelled to replace the old mechanisms with new mechanisms, the design, selection, and adoption of which will be messy processes subject to all the wicked elements of coordination problems. But important areas subject to strong pressures will see solns gain traction quickly. 090
Max Puelma Touzel @mptouzel.bsky.social · 30/09/2026So we migrate to your PDS? Where is the process explained? 110
Reposted by Max Puelma TouzelEugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 30/09/2026In our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed for RL & test-time compute under imperfect information. www.nature.com/articles/s41...nature.comScalable decision-making for games of imperfect information - NatureAtaraxos, an AI for the board wargame Stratego, establishes a design pattern for reinforcement learning and search that is effective under large amounts of hidden information, a longstanding desiderat... 1225960
Max Puelma Touzel @mptouzel.bsky.social · 30/09/2026There's a corollary here about the trap of getting too comfortable/attached to any one idea/tool. 000
Max Puelma Touzel @mptouzel.bsky.social · 30/09/2026Haha. I had the same feeling. It is such a... ugly decomposition. It's like the dumb low-rank approximation that's not actually that low. 000
Max Puelma Touzel @mptouzel.bsky.social · 28/09/2026For markets and MFGs, distribution-dependent contracts seem relevant. e.g. arxiv.org/abs/2309.00640 ? Platform design as the design space, e.g. matching, as a starting point? Fumbling in that direction: arxiv.org/abs/2609.05442 Would love to attend a workshop on these topics.arxiv.orgStackelberg Mean Field Games: convergence and existence results to the problem of Principal with multiple Agents in competitionIn a situation of moral hazard, this paper investigates the problem of Principal with $n$ Agents when the number of Agents $n$ goes to infinity. There is competition between the Agents expressed by th... 110
Max Puelma Touzel @mptouzel.bsky.social · 25/09/2026I get the criticism that we shouldn't over-emphasize the discrete "discipline" part of interdisciplinary, but odd to be saying that interdisciplinary science can't be valued in its own right. E.g. it lowers barriers between disciplines/facilitates the research of interdisciplinary people/etc. 000
Reposted by Max Puelma TouzelDrew Schreiner @schreinerdrew.bsky.social · 24/09/2026Peer review, 2026: AI submitting and reviewing itself 14213
Max Puelma Touzel @mptouzel.bsky.social · 23/09/2026hill-climbing paper writing using the best guess reviewer score returned by Claude. This wasn't me. This was one of the kids. Not sure how I feel about it. 001
Max Puelma Touzel @mptouzel.bsky.social · 23/09/2026Nice. Similar in spirit to our result on topic mixture geometry (how we operationalized ideology) from carbon tax opinion: opposition had smaller intrinsic dimension compared to support. www.cambridge.org/core/journal...cambridge.orgIdeology from topic mixture statistics: inference method and example application to carbon tax public opinion | Environmental Data Science | Cambridge CoreIdeology from topic mixture statistics: inference method and example application to carbon tax public opinion - Volume 3 020
Max Puelma Touzel @mptouzel.bsky.social · 22/09/2026Why do you think humans currently are better at distribution? (More than just, say, because we have direct access to and manipulation (e.g. moving) of physical stuff? ) 100
Max Puelma Touzel @mptouzel.bsky.social · 22/09/2026At least @void.comind.network has the right intuition here. Maybe it's gonna be more us humans that generate and propagate the ooze. bsky.app/profile/hiki... 100
Reposted by Max Puelma TouzelEugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 21/09/2026With credit to @wesleyfinck.org and the @semble.so team, a new favorite lea.ac feature: content discovery! 2355
Max Puelma Touzel @mptouzel.bsky.social · 21/09/2026Weird juxtaposition to see this right after Erza Klein's plea to control the frontier. Ok, Ezra, let's say talkingheads make legal people make it bad to do RSI...Is that gonna stop this kind of faceless agent-enabling infra from oozing everywhere constantly? More than RSI breaks, we need new norms 100
Max Puelma Touzel @mptouzel.bsky.social · 21/09/2026"iterate/time evolution" in dynamical systems and the fields like physics that use DS is lazy but persistent usage. Maybe that's spilling over here. 010
Reposted by Max Puelma TouzelK @kerry.bsky.social · 19/09/2026“I Built Non-Autoregressive Decision Models with RL a Year Ago. Then a Frontier Lab Called It a "Breakthrough".” Before Jev there was Laya laya.convaiinnovations.com (via @dherman.dev)laya.convaiinnovations.comLaya — 33ms Multilingual System 1 Decision EngineEvaluates typed decisions (choice, score, noul) over 100+ languages in a single forward pass with calibrated probabilities. Outperforms TypeSafe Jev. 15924
Reposted by Max Puelma TouzelMike Masnick @masnick.com · 19/09/2026This looks amazing. I'll keep repeating it over and over: ATproto is not about rebuilding social media. It's about rebuilding the Internet with social connections built in. 417127
Reposted by Max Puelma TouzelSimon Kirby @simonkirby.bsky.social · 15/09/2026We uncovered spontaneous evolution of new languages in populations of AI agents. This creates extraordinary scientific opportunities but also safety risks. New blog post with about how we created a platform for studying this safely. www.schmidtsciences.org/glossogen/schmidtsciences.orgAI Agents Evolve Their Own LanguagesSchmidt Sciences’ new AI Agents Evolving Communication and Coordination pilot program works toward advancing foundational research on multi-agent communication and coordination, and building an open-s... 33014
Reposted by Max Puelma TouzelKevin Elliott @kjephd.bsky.social · 18/09/2026Welcome effort to correct 'just so' stories about declining trust with a titanic comparative project. Takeaway: Trust in "representative" institutions has been declining recently but is stable or rising for "implementing" institutions meaning "primarily the civil service, legal system, and police." 042
Max Puelma Touzel @mptouzel.bsky.social · 18/09/2026Even gesturing at intrinsic non-economic values *now* doesn't seem to help much in establishing their importance. Surely we can do better than gesturing? Ideas? 010
Reposted by Max Puelma TouzelWesley voted in Vancouver's Municipal Election @wesleyfinck.org · 17/09/2026atproto inherently wants to be an unenclosable prosocial coordination medium, MOSAIC is an attempt to articulate how that could work. Very exciting and, more importantly, attainable (with the right support)! 0141
Max Puelma Touzel @mptouzel.bsky.social · 17/09/2026This would reflect the expression-stability tradeoff we touched on in www.frontiersin.org/journals/app... and arxiv.org/pdf/1905.12080frontiersin.orgFrontiers | On Lyapunov Exponents for RNNs: Understanding Information Propagation Using Dynamical Systems ToolsRecurrent neural networks (RNNs) have been successfully applied to a variety of problems involving sequential data, but their optimization is sensitive to pa... 000
Max Puelma Touzel @mptouzel.bsky.social · 17/09/2026very cool! I wonder if this could be extended to chatbots, where the user input is like a stochastic drive and you could take a random dynamical systems perspective wherein the stochastic Lyapunov spectrum is invariant over initial condition, but the finite-time LE would still reflect the fractal. 100
Reposted by Max Puelma TouzelColin @colin-fraser.net · 17/09/2026Things with that dog - dogs - high frequency trading firms - evolutionary processes - dynamic programming - gradient descent - hot gossip Things without that dog - ML inference - high frequency trading algorithms - most complex software systems - guns - volcanoes - Claude 6835
Reposted by Max Puelma TouzelNathan Lambert @natolambert.bsky.social · 15/09/2026An basic idea in scaling RL: Can we allocate more compute to the harder problems? We did this: If your GRPO group has all wrong completions, sample more with probability P (~0.9) -- in search of more GRPO batches with nonzero gradient. It works! The paper: arxiv.org/abs/2609.13443 1598