Reposted by Kale-ab TesseraRL & Agents Reading Group @rl-agents-rg.bsky.social · 23/09/2026We are back with the RL & Agents Reading Group tomorrow (Thursday) 3pm, kicking off a three-talk RLC series! Leo Hinckeldey is presenting Assistax, a benchmark for multi-agent assistive robotics. Interested in MARL, ad hoc teamwork, embodied RL? Join us! Paper: arxiv.org/abs/2507.21638 032
Kale-ab Tessera @kale-ab.bsky.social · 13/08/2026I’ll be @ RLC in Montreal 🇨🇦 Sat: Talk on Alem, our benchmark for long-horizon, open-ended coordination @rlvgworkshop.bsky.social (11:15). 👇 Tue: Joining Marcel Hedman who led CODA, a diffusion approach for offline MARL. Hmu to chat multi-agent systems, post-training & open-endedness/RSI 🎉 031
Reposted by Kale-ab TesseraLukas Schäfer @lukaschaefer.bsky.social · 14/07/2026This looks awesome! We need benchmarks that consider coordination between agents. Also super cool to see cross comparison between LLM/ VLM and MARL agents! Congrats to @kale-ab.bsky.social & co. 👏 094
Kale-ab Tessera @kale-ab.bsky.social · 14/07/2026Very cool, especially the struct types + MCP server (when you asked Gemini to take over in the crossword game :0 )! 🔥 050
Reposted by Kale-ab TesseraMarc Lanctot @sharky6000.bsky.social · 14/07/2026I'm delighted to announce the release of OpenSpiel 2.0! ♟️🎲♦️🎉 Structured types for states, observations, and actions, standard trajectories (based on JSON), 19 new games, AlphaZero ported to JAX, Windows PyPI support, language model fine-tuning examples and an MCP server (demo below 🤩👇)! 🧵 1/N 46713
Kale-ab Tessera @kale-ab.bsky.social · 14/07/2026Coordination could be the next frontier for agents. We hope Alem helps drive progress across LLM and MARL systems. 🚀 Huge thanks to Andras Szecsenyi, Cameron Barker, Alex Rutherford, Davide Paglieri, Aidan Scannell, Henry Gouk, Elliot Crowley, Tim Rocktäschel, and Amos Storkey. 030
Kale-ab Tessera @kale-ab.bsky.social · 14/07/2026Links: 📄 Paper: arxiv.org/abs/2606.08340 🌐 Website: alem-world.github.io 💻 Env + harness: github.com/alem-world/a... 🤗 MARL baselines: huggingface.co/alem-world/a... Try your own agent and submit it!arxiv.orgBenchmarking Open-Ended Multi-Agent Coordination in Language AgentsAs language models are increasingly deployed as autonomous agents, they must coordinate with others over long horizons in open-ended interactive tasks. Yet existing evaluations rarely test these deman... 140
Kale-ab Tessera @kale-ab.bsky.social · 14/07/2026There's much more in the paper and website, including coordination by task type, model-size effects, and sampled interactive traces from all 13 models. The agent conversations alone are worth exploring, from role assignment to coordinated revives: alem-world.github.io/traces.html 120
Kale-ab Tessera @kale-ab.bsky.social · 14/07/2026Can stronger models carry weaker teammates in heterogeneous teams? In our initial experiments, not really. Mixed teams perform almost exactly at the average of their homogeneous baselines (+0.1 / −0.1 Total%), neither collapsing to the weakest model nor matching the strongest. 130
Kale-ab Tessera @kale-ab.bsky.social · 14/07/2026What drives coordination? Communication matters most. Removing it drops Gemini’s Coord.% from 17.5 to 5.3 and Gemma-4-31B’s from 8.8 to 3.8, while Base% changes much less. One clue as to why: ~44% of messages address a specific teammate, often to assign work, share intent, or align timing. 151
Kale-ab Tessera @kale-ab.bsky.social · 14/07/2026Looking beyond aggregate scores: 1. Base competence ≠ coordination: smaller open-weight models beat GPT-5.4 on coordination in some settings, with non-overlapping CIs. 2. Survival ≠ progress: MARL matches the strongest LLM's return in far fewer steps, echoing ARC-AGI-3 focus on action efficiency. 130
Kale-ab Tessera @kale-ab.bsky.social · 14/07/2026On the Hard setting, zero-shot Gemini 3.1 Pro performs comparably to the best MARL agent trained for 1 billion steps (17.5 vs 17.6 Coord.%, 15.4 vs 15.3 Total.%). *(Different interfaces, same underlying environment). 130
Kale-ab Tessera @kale-ab.bsky.social · 14/07/2026Current LLMs are far from solving Alem. Across 13 LLMs, average performance is only ~6% normalised return. Even the strongest MARL baseline trained for 3 billion steps reaches just 18% of max total reward. But one result surprised us: 120
Kale-ab Tessera @kale-ab.bsky.social · 14/07/2026Agents navigate 9 procedural levels, pursue 93 achievements, and act for up to 10,000 steps. Alem generates not just the world, but the coordination problem itself, with controllable difficulty. Text, pixel, and symbolic interfaces let LLMs, VLMs, humans, and MARL agents face the same environment. 140
Kale-ab Tessera @kale-ab.bsky.social · 14/07/2026LLM agents are improving at long-horizon tasks. But deployed agents will also need to coordinate with each other. Yet most benchmarks still test agents alone or in short, structured collaboration. So we built Alem ("world" in Amharic), a JAX benchmark built on Craftax-Coop/Craftax. 140
Kale-ab Tessera @kale-ab.bsky.social · 14/07/2026Can LLM agents coordinate in long-horizon, open-ended worlds? We test 13 LLMs in Alem, a new benchmark. Most struggle, averaging ~6% normalised return. Yet on Hard, zero-shot Gemini 3.1 Pro matches the best MARL agent after 1B training steps. Our ablations show communication matters most. 🧵 13810
Kale-ab Tessera @kale-ab.bsky.social · 29/05/20262. Def agreed! I think our metrics would still be somewhat useful in mixed-motive -- since they look at partial obs & coordination/inf, and have no real links to reward structure. You would def need new metrics to capture the essence of mixed motive, though. A promising future direction 😄 020
Kale-ab Tessera @kale-ab.bsky.social · 29/05/2026GDI is much cleaner & robust than plain directed info, we really wanted to use it but ran out of time before the deadline! 😔 120
Kale-ab Tessera @kale-ab.bsky.social · 29/05/2026From my view, the plasticity paper is about how actions shape future obs and obs shape future actions, using their GDI metric. Most of ours is about influence between agents, but our obs → action metric could be seen on a similar "plasticity" axis. @dabelcs.bsky.social Pls correct me if I am wrong! 130
Kale-ab Tessera @kale-ab.bsky.social · 29/05/2026Thanks so much for the kind words Marc, it really means a lot, especially coming from you! 🙌 1. Awesome question & huge fan of that paper! Dave is on our paper too and was critical to the inspiration & framing of our work. :) ... 120
Kale-ab Tessera @kale-ab.bsky.social · 26/05/2026Presenting this tomorrow at 14:00–15:45, LEARN 2 Session, Akamas Room. Ping me if you are around at AAMAS & want to chat! 👐 Slides: kaleabtessera.com/assets/slide...kaleabtessera.com 030
Kale-ab Tessera @kale-ab.bsky.social · 25/05/2026We open-sourced the library for computing the metrics and have a fun website you can play around with to get the intuition. 📜: arxiv.org/abs/2602.20804 💻: kaleabtessera.com/probing 🧑💻: github.com/KaleabTesser...arxiv.orgProbing Dec-POMDP Reasoning in Cooperative MARLCooperative multi-agent reinforcement learning (MARL) is typically framed as a decentralised partially observable Markov decision process (Dec-POMDP), a setting whose hardness stems from two key chall... 130
Kale-ab Tessera @kale-ab.bsky.social · 25/05/2026We hope these tools can be useful in building the next generation of multi-agent envs, where partial observability and coordination are non-optional! 🚀 Huge thanks to my wonderful collaborators - @leohinck.bsky.social , @ricczamboni.bsky.social , @dabelcs.bsky.social , Amos Storkey. 🎉 130
Kale-ab Tessera @kale-ab.bsky.social · 25/05/2026Few benchmarks jointly test partial observability and coordination ⁉️- MPE was the only suite where every scenario satisfied all diagnostic criteria. Many detected dependencies are above null baselines but still modest, leaving plenty of room to design better envs. 🔥 120
Kale-ab Tessera @kale-ab.bsky.social · 25/05/2026We use these probes across 37 scenarios and 7 envs. Three takeaways: 𝟭) history dependence != history utility (only 43% need memory for high returns); 𝟮) hidden state and teammate info are separate difficulty drivers; 𝟯) sync and temporal coordination differ across envs. 120
Kale-ab Tessera @kale-ab.bsky.social · 25/05/2026🪄To study this, we introduce information-theoretic probes for behaviours that the envs induce under standard MARL algorithms (IPPO/MAPPO). They measure how actions depend on obs, histories, teammate-private information, and other agents' actions, beyond just returns. 130
Kale-ab Tessera @kale-ab.bsky.social · 25/05/2026As we build more cooperative multi-agent envs, we should ask: are we really testing the properties that make Dec-POMDPs hard, partial observability and decentralised coordination, or can agents succeed via shortcuts? 🤔 Our #AAMAS2026 Oral probes this question 🔎🧵 3165
Reposted by Kale-ab TesseraRiccardo Zamboni @ricczamboni.bsky.social · 21/05/2026Next week I will be at AAMAS looking in awe @kale-ab.bsky.social presenting as an oral our paper on probing actual reasoning capabilities in Cooperative Multi-Agent Reinforcement Learning. More than happy to chat about this project or research in general! 142
Reposted by Kale-ab TesseraRL & Agents Reading Group @rl-agents-rg.bsky.social · 21/01/2026Reading group TOMORROW 3-4pm UK! Joe Marino (Google DeepMind) will present SIMA 2, a generalist embodied agent designed to operate across a wide range of 3D virtual worlds 🌎 Join here: edinburgh-rl.github.io/reading-group 074
Reposted by Kale-ab TesseraRL & Agents Reading Group @rl-agents-rg.bsky.social · 09/01/2026Reading group TODAY 2pm BST! "Despite years of research in offline reinforcement learning, the field has failed to deliver major breakthroughs..." Matthew Jackson and Jarek Liesen (Oxford) will present Unifloral - unified implementations and evaluations for offline RL: arxiv.org/abs/2504.11453 031
Kale-ab Tessera @kale-ab.bsky.social · 04/12/2025Happening now - Exhibit Hall C,D,E poster #404 I heard there will be good vibes at this poster 🤙 031
Kale-ab Tessera @kale-ab.bsky.social · 29/11/2025First time in a Waymo. Honestly, a pretty surreal experience! Surprised by how smooth the ride was and how quickly I felt comfortable in the car 😮 050
Kale-ab Tessera @kale-ab.bsky.social · 26/11/2025If you are around and want to chat about multi-agent systems (MARL, agentic systems), open-endedness, environments, or anything related, please let me know! 🎉 060
Kale-ab Tessera @kale-ab.bsky.social · 26/11/2025Thrilled to present HyperMARL at #NeurIPS2025 in San Diego next week! 🚀 (Amos will present at @euripsconf.bsky.social too.) TL;DR: Coupling obs and agent IDs can hurt performance in MARL. Agent-conditioned hypernets cleanly decouple grads and enable specialisation. 📜: arxiv.org/abs/2412.04233 3135
Kale-ab Tessera @kale-ab.bsky.social · 16/11/2025I think most people judge reputation from high-level things e.g. num of accepted papers, and very few people actually read these papers. This means you can game the system with LLM generated papers with little consequences, and this makes things frustrating for everyone. 000
Reposted by Kale-ab TesseraRL & Agents Reading Group @rl-agents-rg.bsky.social · 06/11/2025Reading group today at 2pm BST! We are starting our NeurIPS series with Sable and Oryx, sequence models for scalable multi-agent coordination from the RL Research Team at InstaDeep. 🚀 Papers: - Sable: bit.ly/3Lme7jH - Oryx: bit.ly/47GJb4T Meeting: - bit.ly/3JoEbtU 031
Reposted by Kale-ab TesseraRL & Agents Reading Group @rl-agents-rg.bsky.social · 03/09/2025📢 RL reading group Thursday @ 16:00 BST 📢 Speaker: Alex Lewandowski Title: The World Is Bigger: A Computationally-Embedded Perspective on the Big World Hypothesis 🌍 Details: edinburgh-rl.github.io/reading-groupedinburgh-rl.github.ioUoE RL Reading GroupUniversity of Edinburgh Reinforcement Learning Reading Group 063
Kale-ab Tessera @kale-ab.bsky.social · 19/08/2025Refreshing to see posts like this compared to "we have 15 papers accepted at X" 🙌 110
Reposted by Kale-ab TesseraEugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 19/08/2025None of our impactful papers have had an easy path through traditional venues. Most cited paper? Rejected four times. Most impactful paper? Poster at a conference. But none of it matters because arxiv makes everything work 61076
Kale-ab Tessera @kale-ab.bsky.social · 18/08/2025Great first couple of days at DLI @deeplearningindaba.bsky.social in Kigali 🇷🇼, some highlights include amazing talks talks by @verenarieser.bsky.social and Max Welling, great pracs and tuts, and of course the opening party ( before the rain 😢) 🎉 #DLI2025 041
Reposted by Kale-ab TesseraDeep Learning Indaba @deeplearningindaba.bsky.social · 17/08/2025We’re excited to unveil the first #DLI2025 lineup of tutorials and practicals: ✨ Machine Learning Foundations ✨ Generative Models & LLMs for African languages All tutorial content will also be available online after the Indaba. Don’t miss out, subscribe here 👉 lnkd.in/eCgXRqsV 023
Kale-ab Tessera @kale-ab.bsky.social · 03/08/2025🇨🇦 Heading to @rl-conference.bsky.social next week to present HyperMARL (@cocomarl-workshop.bsky.social) and Remember Markov (Finding The Frame Workshop). If you are around, hmu, happy to chat about Multi-Agent Systems (MARL, agentic systems), open-endedness, environments, or anything related! 🎉 092
Reposted by Kale-ab TesseraDeep Learning Indaba @deeplearningindaba.bsky.social · 30/07/2025We are thrilled to announce our next keynote speaker @wellingmax.bsky.social, Professor at the University of Amsterdam, Visiting Professor at Caltech and CTO & Co-Founder of CuspAI. Catch his talk “How AI could transform the sciences” on August 18 at 4:30 PM GMT+2. #DLI2025 011
Reposted by Kale-ab TesseraRL & Agents Reading Group @rl-agents-rg.bsky.social · 24/07/2025RL reading group TODAY @ 15:00 BST 🔥 Speaker: Cam Allen (Postdoc, UC Berkeley) Title: The Agent Must Choose the Problem Model Details: edinburgh-rl.github.io/reading-groupedinburgh-rl.github.ioUoE RL Reading GroupUniversity of Edinburgh Reinforcement Learning Reading Group 031
Kale-ab Tessera @kale-ab.bsky.social · 23/07/2025Always nice to see when simpler methods + good evaluations > more complicated ones. 👌 010
Kale-ab Tessera @kale-ab.bsky.social · 10/07/2025Reading group is back for those interested in RL/MARL/agents/open-endedness and alike... First session today at 3pm BST, @mattieml.bsky.social is presenting the Simplifying TD learning/PQN paper. 🎉 Meeting link: bit.ly/4lfdaGR Sign up: bit.ly/40xNQDR 031
Reposted by Kale-ab TesseraRL & Agents Reading Group @rl-agents-rg.bsky.social · 10/07/2025Hello world! This is the RL & Agents Reading Group We organise regular meetings to discuss recent papers in Reinforcement Learning (RL), Multi-Agent RL and related areas (open-ended learning, LLM agents, robotics, etc). Meetings take place online and are open to everyone 😊 13712