Reposted by Manfred DiazMarc Lanctot @sharky6000.bsky.social · 14/07/2026I'm delighted to announce the release of OpenSpiel 2.0! ♟️🎲♦️🎉 Structured types for states, observations, and actions, standard trajectories (based on JSON), 19 new games, AlphaZero ported to JAX, Windows PyPI support, language model fine-tuning examples and an MCP server (demo below 🤩👇)! 🧵 1/N 46713
Reposted by Manfred DiazJoel Z Leibo @jzleibo.bsky.social · 21/03/2026New paper: “𝐀 𝐓𝐡𝐞𝐨𝐫𝐲 𝐨𝐟 𝐀𝐩𝐩𝐫𝐨𝐩𝐫𝐢𝐚𝐭𝐞𝐧𝐞𝐬𝐬 𝐓𝐡𝐚𝐭 𝐀𝐜𝐜𝐨𝐮𝐧𝐭𝐬 𝐟𝐨𝐫 𝐍𝐨𝐫𝐦𝐬 𝐨𝐟 𝐑𝐚𝐭𝐢𝐨𝐧𝐚𝐥𝐢𝐭𝐲” Agent-based models of social order work better when agents act by predictive pattern completion from prefix (culture/context) to suffix (action) than when they act through expected value maximization 43511
Reposted by Manfred DiazMarc Lanctot @sharky6000.bsky.social · 15/01/2026Hello all! 👋 I’m delighted to share a 🚨 new preprint 🚨: “Active Evaluation of General Agents: Problem Definition and Comparison of Baseline Algorithms”. A paper thread! 🤩📄🧵 1/N 25712
Manfred Diaz @manfreddiaz.bsky.social · 29/11/2025Maybe the general intelligence has always been behind the algorithm or the prompt? No publicly available eval seems to be safe from researchers overfitting. 000
Manfred Diaz @manfreddiaz.bsky.social · 08/08/2025@sharky6000.bsky.social this may be of interest! 130
Manfred Diaz @manfreddiaz.bsky.social · 16/06/2025I was following this one during the COVID pandemic, but it has been inactive for quite some time. The original talks' recordings are amazing, though! 110
Manfred Diaz @manfreddiaz.bsky.social · 05/06/2025Yeah, it's been a period for all of us simultaneously! I have also been pretty busy with thesis/job search. Hopefully, it will be back running in the Fall term! 010
Manfred Diaz @manfreddiaz.bsky.social · 23/05/2025@aamasconf.bsky.social 2025 was very special for us! We had the opportunity. to present a tutorial on general evaluation of AI agents, and we got a best paper award! Congrats, @sharky6000.bsky.social and the team! 🎉 0121
Reposted by Manfred DiazMarc Lanctot @sharky6000.bsky.social · 18/05/2025In the afternoon we will be giving a tutorial on general evaluation of AI agents. sites.google.com/view/aamas20... 10/Nsites.google.comA Tutorial on General Evaluation of AI AgentsArtificial Intelligence (AI) and machine learning (ML), in particular, have emerged as scientific disciplines concerned with understanding and building single and multi-agent systems with the ability ... 141
Reposted by Manfred DiazJoel Z Leibo @jzleibo.bsky.social · 09/05/2025Announcing our latest arxiv paper: Societal and technological progress as sewing an ever-growing, ever-changing, patchy, and polychrome quilt arxiv.org/abs/2505.05197 We argue for a view of AI safety centered on preventing disagreement from spiraling into conflict.arxiv.orgSocietal and technological progress as sewing an ever-growing, ever-changing, patchy, and polychrome quiltArtificial Intelligence (AI) systems are increasingly placed in positions where their decisions have real consequences, e.g., moderating online spaces, conducting research, and advising on policy. Ens... 1236
Reposted by Manfred DiazJoel Z Leibo @jzleibo.bsky.social · 22/04/2025First LessWrong post! Inspired by Richard Rorty, we argue for a different view of AI alignment, where the goal is "more like sewing together a very large, elaborate, polychrome quilt", than it is "like getting a clearer vision of something true and deep" www.lesswrong.com/posts/S8KYwt...lesswrong.comSocietal and technological progress as sewing an ever-growing, ever-changing, patchy, and polychrome quilt — LessWrongWe can just drop the axiom of rational convergence. 351
Manfred Diaz @manfreddiaz.bsky.social · 16/04/2025The quality of London's museums is just amazing! Enjoy! 020
Reposted by Manfred DiazJoel Z Leibo @jzleibo.bsky.social · 01/04/2025In case folks are interested, here's a video of a talk I gave at MIT a couple weeks ago: youtu.be/FmN6fRyfcsY?...youtu.beA Theory of Appropriateness with Applications to Generative Artificial IntelligenceYouTube video by MITCBMM 073
Reposted by Manfred DiazMarc Lanctot @sharky6000.bsky.social · 28/03/2025Our new evaluation method, Soft Condorcet Optimization is now available open-source! 👍 Both the sigmoid (smooth Kendall-tau) and Fenchel-Young (perturbed optimizers) versions. Also, an optimized C++ implementation that is ~40X faster than the Python one. 🤩⚡ github.com/google-deepm... 0163
Reposted by Manfred DiazMarc Lanctot @sharky6000.bsky.social · 26/03/2025Working at the intersection of social choice and learning algorithms? Check out the 2nd Workshop on Social Choice and Learning Algorithms (SCaLA) at @ijcai.bsky.social this summer. Submission deadline: May 9th. I attended last year at AAMAS and loved it! 👍 sites.google.com/corp/view/sc...sites.google.comSCaLA-25A workshop connecting research topics in social choice and learning algorithms. 0196
Manfred Diaz @manfreddiaz.bsky.social · 06/03/2025If the AAMAS website is a good reference for this, it may not be, but uncertain atm. 110
Manfred Diaz @manfreddiaz.bsky.social · 04/03/2025Come to understand ML evaluation from first principles! We have put together a great AAMAS tutorial covering statistics, probabilistic models, game theory, and social choice theory. Bonus: a unifying perspective of the problem leveraging decision-theoretic principles! Join us on May 19th! 161
Manfred Diaz @manfreddiaz.bsky.social · 04/03/2025Re #2: The key finding there is that the stationary points of SCO contain the margin matrix but, as I said in the note, there is still more work to do! 110
Manfred Diaz @manfreddiaz.bsky.social · 04/03/2025Thanks! I have been meaning to update the manuscript to standalone without the main paper but instead I may have change the content to a different format 😉. Coming soon! 210
Manfred Diaz @manfreddiaz.bsky.social · 25/02/2025Ah, I see the confusion... I never used the "identically distributed assumption," only the independence assumption (from 8 to 9). 010
Manfred Diaz @manfreddiaz.bsky.social · 25/02/2025I'm not sure if I understood your question correctly, but yes? As the post you shared says, "Voila! We have shown that minimizing the KL divergence amounts to finding the maximum likelihood estimate of θ." Maybe I am missing your point 😬 200
Manfred Diaz @manfreddiaz.bsky.social · 25/02/2025Elo drives most LLM evaluations, but we often overlook its assumptions, benefits, and limitations. While working on SCO, we wanted to understand the SCO-Elo distinction, so I looked and uncovered some intriguing findings and documented them in these notes. I hope you find them valuable! 021
Reposted by Manfred DiazMarc Lanctot @sharky6000.bsky.social · 24/02/2025Looking for a principled evaluation method for ranking of *general* agents or models, i.e. that get evaluated across a myriad of different tasks? I’m delighted to tell you about our new paper, Soft Condorcet Optimization (SCO) for Ranking of General Agents, to be presented at AAMAS 2025! 🧵 1/N 16517
Manfred Diaz @manfreddiaz.bsky.social · 20/02/2025I had the convexity results for the online pairwise update (Section B.1.1.1) in my notes (manfreddiaz.github.io/assets/pdf/s...), but it is not clear to me if they hold for the other non-online settings. Worth taking a more detailed pass over the paper!manfreddiaz.github.io 020
Manfred Diaz @manfreddiaz.bsky.social · 20/02/2025That's a nice finding, @sacha2.bsky.social! @sharky6000.bsky.social I skimmed over it, and it seems neat! There is an important distinction, though. They work with the "online" Elo regime, departing from the traditional gradient/batch gradient descent updates. (e.g., FIDE doesn't use online updates) 120
Manfred Diaz @manfreddiaz.bsky.social · 12/02/2025Not that Michael Jordan, but this one en.wikipedia.org/wiki/Michael...en.wikipedia.orgMichael I. Jordan - Wikipedia 130
Manfred Diaz @manfreddiaz.bsky.social · 11/02/2025I believe this example conveys, as Prof. Jordan hinted, the need for fresh conceptual frameworks that shift our perspective, help us avoid conceptual confusion, and increase our ability to build the future of AI. I believe ML-SoA provides such framework, but I’d love to hear more perspectives! 140
Manfred Diaz @manfreddiaz.bsky.social · 11/02/2025Through the ML-SoA design-centric view, reward hacking, misspecification and similar issues originate from designers' bounded rationality. The surprising “alignment problems” (www.youtube.com/watch?v=tlOI...) reflect ML designers' inability to comprehend the consequences of the programs they write.youtube.comCoastRunners 7YouTube video by Jack Clark 150
Manfred Diaz @manfreddiaz.bsky.social · 11/02/2025For instance, in the AI safety literature, we hear phrases such as “the agent hacked the reward”, “the robot gamed the reward”, and similar statements. But does an ML model really "hack" or "game" anything? Or is this just a misleading metaphor? Chances are, it’s the latter. 120
Manfred Diaz @manfreddiaz.bsky.social · 11/02/2025None of these benefits is more critical than removing, to a greater extent, the anthropomorphism that has existed in AI, currently exacerbated by the rise of LLMs, whose principal peril is the conceptual confusion it sometimes causes. Let me clarify what I mean by "conceptual confusion". 110
Manfred Diaz @manfreddiaz.bsky.social · 11/02/2025We recently argued here (openreview.net/forum?id=LNY...) that examining ML through this perspective, which I call Machine Learning Through the Sciences of the Artificial (ML-SoA), offers solid foundations for understanding AI and ML in a broader context and introduces multiple benefits.openreview.netMilnor-Myerson Games and The Principles of Artificial...In this paper, we introduce Milnor-Myerson games, a multiplayer interaction structure at the core of machine learning (ML), to shed light on the fundamental principles and implications the... 130
Manfred Diaz @manfreddiaz.bsky.social · 11/02/2025Back in the 1960, Herbert Simon defined in "The Sciences of the Artificial" an artificial entity as a product of human creation. In this context, AI shares foundations with other scientific theories that study human-designed structures such as institutions, markets, economies, and elections. 130
Manfred Diaz @manfreddiaz.bsky.social · 11/02/2025Last week, Michael I. Jordan's insightful talk at the AI Action Summit (www.youtube.com/live/W0QLq4q...) reminded us of the meaningful connections between AI, ML, economics, game theory, and mechanism design. But I'd argue the relationship goes deeper—it's profound, historical, and foundational. ⬇️youtube.comAI, Science and Society Conference - AI ACTION SUMMIT - DAY 1YouTube video by IP Paris 180
Manfred Diaz @manfreddiaz.bsky.social · 11/02/2025Indeed! Some of the effects you predicted in your post are already being measured on StackOverflow (arxiv.org/abs/2307.07367).arxiv.orgAre Large Language Models a Threat to Digital Public Goods? Evidence from Activity on Stack OverflowLarge language models like ChatGPT efficiently provide users with information about various topics, presenting a potential substitute for searching the web and asking people for help online. But since... 281
Manfred Diaz @manfreddiaz.bsky.social · 08/02/2025I completely agree! Decision theory, game theory, and mechanism design offer better foundations for understanding AI and ML in a broader context. Framing ML this way has many benefits, and we recently argued for this position (openreview.net/forum?id=LNY...).openreview.netMilnor-Myerson Games and The Principles of Artificial...In this paper, we introduce Milnor-Myerson games, a multiplayer interaction structure at the core of machine learning (ML), to shed light on the fundamental principles and implications the... 041
Manfred Diaz @manfreddiaz.bsky.social · 07/01/2025I wonder what the source of those trending topics are. Is it the whole platform? The people you follow? Content in the Discover tab? 🤷🏼♂️ 050
Manfred Diaz @manfreddiaz.bsky.social · 07/01/2025The Discover tab shows some kind of trending topics every time. For instance, I learned about the possibility of the Prime Minister resigning here on BlueSky yesterday. 170
Manfred Diaz @manfreddiaz.bsky.social · 01/01/2025I already watched Slow Horses, one of my favourites! And thanks to your recommendation, I'm now a fan of Silo! 040
Reposted by Manfred DiazJoel Z Leibo @jzleibo.bsky.social · 31/12/2024Very happy to announce the publication of our latest paper: A theory of appropriateness with applications to generative artificial intelligence arxiv.org/abs/2412.19010 And happy new year everyone!arxiv.orgA theory of appropriateness with applications to generative artificial intelligenceWhat is appropriateness? Humans navigate a multi-scale mosaic of interlocking notions of what is appropriate for different situations. We act one way with our friends, another with our family, and yet... 2317
Manfred Diaz @manfreddiaz.bsky.social · 10/12/2024NeurIPS generally both livestreams and records most if not all sessions and make them available after the conference. Takes a while but we eventually get them! 020
Manfred Diaz @manfreddiaz.bsky.social · 10/12/2024I didn't know about this! Sad that I'll miss it but I'll definitely watch the recording! 🙂 120
Manfred Diaz @manfreddiaz.bsky.social · 04/12/2024not much I can say atm 🙂! But we are using a bit of social choice theory to propose an alternative to alignment. It should be coming out soon! 110
Manfred Diaz @manfreddiaz.bsky.social · 04/12/2024Interesting! I'd definitely take a look! Thank you both! 120