Reposted by Mirco MuttiClément Canonne @ccanonne.github.io · 01/08/2026"We're all worried," as what it means to do research (in my field, Theoretical CS) seems to be shifting, and shifting fast. What to do? Senior researchers must lead by example, knowing that not everything will pan out. What I'm suggesting below may not work everywhere, but here's my own advice: 1/ 620351
Reposted by Mirco MuttiRiccardo Zamboni @ricczamboni.bsky.social · 02/04/2026🚨🚨🚨 Later today I am going to present at @rl-agents-rg.bsky.social’s reading group one research line with @mircomutti.bsky.social that I’m really excited about: behavior compression via unsupervised RL! 062
Reposted by Mirco MuttiELLIS @ellis.eu · 04/03/2026📣 Reinforcement Learning Summer School is returning to Milan in 2026! Co-organized with @ellisunitmilan.bsky.social & designed for Master's and PhD students on RL theory, multi-agent systems, RL & LLMs, real-world applications... 📍 Milan 🇮🇹 📅 3-12 June ⏰ Apply by 27 March 🔗 bit.ly/4b2Plhp 02411
Mirco Mutti @mircomutti.bsky.social · 16/12/2025for inverse: I like a lot the conceptualization of the problem in the works by Alberto & Filippo, such as - proceedings.mlr.press/v202/metelli... - arxiv.org/pdf/2501.07996 (may be biased here bc I collaborated in some of those)proceedings.mlr.press 010
Mirco Mutti @mircomutti.bsky.social · 16/12/2025for imitation: - arxiv.org/pdf/2503.09722 around "separation between bc in discrete and continuous settings" and followups - dylanfoster.net/il-tutorial/ tutorial on foundations of imitation learning by Max, Dylan, Adam may also be a useful lookup 100
Mirco Mutti @mircomutti.bsky.social · 28/11/2025Absolutely, come to the poster! Some say Riccardo's aura will be hovering around 010
Mirco Mutti @mircomutti.bsky.social · 18/11/2025No, but since the pc explicitly suggested to post on the 20th, I think most people will comply 100
Reposted by Mirco MuttiTransactions on Machine Learning Research @tmlrorg.bsky.social · 14/10/2025As Transactions on Machine Learning Research (TMLR) grows in number of submissions, we are looking for more reviewers and action editors. Please sign up! Only one paper to review at a time and <= 6 per year, reviewers report greater satisfaction than reviewing for conferences! 11112
Reposted by Mirco MuttiEWRL @ewrl-org.bsky.social · 13/08/2025📣Registration for EWRL is now open📣 Register now 👇 and join us in Tübingen for 3 days (17th-19th September) full of inspiring talks, posters and many social activities to push the boundaries of the RL community!site.pheedloop.comPheedLoopPheedLoop: Hybrid, In-Person & Virtual Event Software 084
Mirco Mutti @mircomutti.bsky.social · 29/07/2025Looks interesting, but cannot access the url or find the report anywhere 000
Mirco Mutti @mircomutti.bsky.social · 24/07/2025That’s my little #ICML2025 convex RL roundup! If you know of other cool work in this space (or are working on one), feel free to reply and share. Hope to see even more work on convex RL variations 🚀 n/n 010
Mirco Mutti @mircomutti.bsky.social · 24/07/2025📄Flow density control – @desariky.bsky.social et al Bridging convex RL with generative models: How to steer diffusion/flow models to optimize non-linear user-specified utilities (beyond just entropy reg fine tuning)? 📍 EXAIT workshop 🔗 openreview.net/pdf?id=zOgAx... 7/n 120
Mirco Mutti @mircomutti.bsky.social · 24/07/2025📄Towards unsupervised multi-agent RL – @ricczamboni.bsky.social et al (yours truly!) Still in the convex Markov games space—this work explores more tractable objectives for the learning setting. 📍EXAIT workshop 🔗https://openreview.net/pdf?id=A1518D1Pp9 6/n 100
Mirco Mutti @mircomutti.bsky.social · 24/07/2025📄Convex Markov games – Ian Gemp et al If you can 'convexify' MDPs, so you can do for Markov games. These two papers lay out a general framework + algorithms for the zero-sum version. 🔗https://openreview.net/pdf?id=yIfCq03hsM 🔗https://openreview.net/pdf?id=dSJo5X56KQ 5/n 110
Mirco Mutti @mircomutti.bsky.social · 24/07/2025📄The number of trials matters in infinite-horizon MDPs – @pedrosantospps.bsky.social et al A deeper look at how the number of realizations used to compute F affects the convex RL problem in infinite horizon settings. 🔗https://openreview.net/pdf?id=I4jNAbqHnM 4/n 110
Mirco Mutti @mircomutti.bsky.social · 24/07/2025📄Online episodic convex RL – Bianca Marni Moreno et al Regret bounds for online convex RL, where F^t is adversarial and revealed only after each episode (or just evaluated on the given trajectory in a bandit feedback variation) 🔗https://openreview.net/pdf?id=d8xnwqslqq 3/n 100
Mirco Mutti @mircomutti.bsky.social · 24/07/2025🔍 Convex RL Standard RL optimizes a linear objective: ⟨d^π, r⟩. Convex RL generalizes this to any F(d^π), where F is non-linear (originally assumed convex—hence the name). This framework subsumes: • Imitation • Risk sensitivity • State coverage • RLHF ...and more. 2/n 100
Mirco Mutti @mircomutti.bsky.social · 24/07/2025Walking around posters at @icmlconf.bsky.social, I was happy to see some buzz around convex RL—a topic I’ve worked on and strongly believe in. Thought I’d share a few ICML papers on this direction. Let’s dive in👇 But first… what is convex RL? 🧵 1/n 151
Mirco Mutti @mircomutti.bsky.social · 15/07/2025To learn more: - come at our poster (n. 908) on Thursday morning session #ICML2025 - read the preprint arxiv.org/abs/2504.04505 - watch the seminar youtube.com/watch?v=pNos... n/narxiv.orgA Classification View on Meta Learning BanditsContextual multi-armed bandits are a popular choice to model sequential decision-making. E.g., in a healthcare application we may perform various tests to asses a patient condition (exploration) and t... 010
Mirco Mutti @mircomutti.bsky.social · 15/07/2025This is how we got "A classification view on meta learning bandits", a joint work with awesome collaborators Jeongyeol, Shie, and @aviv-tamar.bsky.social 7/n 110
Mirco Mutti @mircomutti.bsky.social · 15/07/2025The regret bounds depend on an instance-dependent "classification coefficient", which suggests classification really captures the complexity of the problem rather than being a mere implementation tool 6/n 110
Mirco Mutti @mircomutti.bsky.social · 15/07/2025For the latter, we show exploration is *interpretable*, as it is implemented by a shallow decision tree of simple constant action policies, and *efficient*, giving upper/lower bounds to the regret 5/n 110
Mirco Mutti @mircomutti.bsky.social · 15/07/2025Yes, apparently! A simple algorithm that classifies the latent (condition) with a decision tree (img above right) and then exploits the best action for the classified latent does the job 4/n 110
Mirco Mutti @mircomutti.bsky.social · 15/07/2025Humans typically develop a standard strategy prescribing a sequence of tests to diagnose the condition before committing to the best treatment (see img left). Can we design a bandit algorithm that learns a similarly interpretable exploration but it's also provably efficient? 3/n 110
Mirco Mutti @mircomutti.bsky.social · 15/07/2025Think about a setting in which we aim to converge on the best treatment (action) for a given patient (context) with some unknown condition (latent). The difference between how humans and bandits approach this same problem is striking: 2/n 110
Mirco Mutti @mircomutti.bsky.social · 15/07/2025Would you trust a bandit algorithm to make decisions on your health or investments? Common exploration mechanisms are efficient but scary. In our latest work at @icmlconf.bsky.social, we reimagine bandit algorithms to get *efficient* and *interpretable* exploration. A 🧵 below 1/n 130
Mirco Mutti @mircomutti.bsky.social · 09/07/2025Here we have an original take on how to make the best of parallel data collection for RL. Don't miss the poster at ICML, we're curious to hear what y'all think! Kudos to the awesome students Vincenzo and @ricczamboni.bsky.social for their work under the wise supervision of Marcello. 010
Reposted by Mirco MuttiAmir-massoud Farahmand @sologen.bsky.social · 09/07/2025What do we talk about when we talk about the Bellman Optimality Equation? If we think carefully, we are (implicitly) making three claims. #FoundationsOfReinforcementLearning #sneakpeek 061
Reposted by Mirco MuttiGautam Kamath @gautamkamath.com · 07/07/2025System is so broken: - researchers write papers no one reads - reviewers don't have time to review, shamed to coauthors, use LLMs instead of reading - authors try to fool said LLMs with prompt injection - evaling researchers based on # of papers (no time to read) Dystopic. 1010610
Reposted by Mirco MuttiEWRL @ewrl-org.bsky.social · 08/04/2025Mark your calendars, EWRL is coming to Tübingen! 📅 When? September 17-19, 2025. More news to come soon, stay tuned! 03714
Reposted by Mirco MuttiTim van Erven @timvanerven.nl · 08/04/2025Just enjoyed @mircomutti.bsky.social's seminar talk about interpretable meta-learning of contextual bandit types. The recording is available in case you missed it: youtu.be/pNos7AHGMXwyoutu.beTheory of Interpretable AI Seminar: Mirco MuttiYouTube video by Theory of Interpretable AI Seminar 081
Mirco Mutti @mircomutti.bsky.social · 08/04/2025Happening today! Join us if you want to hear about our take on interpretable exploration for multi-armed bandits. If interested but cannot join, here's the arxiv arxiv.org/abs/2504.04505 Joint work with Jeongyeol, Shie, and @aviv-tamar.bsky.social 001
Reposted by Mirco MuttiTim van Erven @timvanerven.nl · 27/03/2025⏰⏰Theory of Interpretable AI Seminar ⏰⏰ In two weeks, April 8, Mirco Mutti will talk about "A Classification View on Meta Learning Bandits" 063
Mirco Mutti @mircomutti.bsky.social · 17/03/2025The right review form is: - Summary - Comment - Evaluation Curious of alternative arguments, as it looks like conferences are going in a different direction 020
Mirco Mutti @mircomutti.bsky.social · 20/02/2025Awesome! Have a look at this thread to see some nice multi-object manipulation results 020
Reposted by Mirco MuttiAldo Pacchiano @aldopacchiano.bsky.social · 30/01/2025[4/5] “A Theoretical Framework for Partially-Observed Reward States in RLHF” develops and analyzes a model for RLHF where we posit the human feedback to be generated by a stateful labeler. @mircomutti.bsky.social 141
Mirco Mutti @mircomutti.bsky.social · 13/12/2024Among other treats, they'll show you how the common notion of feasible reward set is not suitable here (even in linear MDPs). Enters *reward compatibility*: A new theory-backed solution concept that allows to rephrase inverse RL into a tractable classification task 000
Mirco Mutti @mircomutti.bsky.social · 13/12/2024If interested on our take on addressing inverse RL in large state spaces, go to meet @filippo_lazzati and @alberto_metelli in the poster session 5 #NeurIPS2024 today (paper -> arxiv.org/abs/2406.03812) 152
Reposted by Mirco MuttiAndrea Celli @acelli.bsky.social · 28/11/2024I will soon be opening a call for a postdoctoral position in online learning and algorithmic game theory, starting in 2025, funded by my ERC at Bocconi University. If you're interested, feel free to reach out. If you're not personally interested but know someone who might be, please let them know! 184
Reposted by Mirco MuttiDr. Angelica Lim @petitegeek.bsky.social · 23/11/2024This is nice brain candy for the affective computing crowd 011
Mirco Mutti @mircomutti.bsky.social · 22/11/2024These two give (mostly orthogonal) perspectives on modelling evolving "internal states" of the human evaluator while interacting with the system arxiv.org/pdf/2402.03282 arxiv.org/pdf/2405.17713 (shameless advertisement alert) 130
Reposted by Mirco MuttiEugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 09/11/2024If you're an RL researcher or RL adjacent, pipe up to make sure I've added you here! go.bsky.app/3WPHcHg 527127
Mirco Mutti @mircomutti.bsky.social · 20/11/2024Hello there! I'm new here and interested in AI -especially reinforcement learning- and keeping up with the latest in research. I'll occasionally share updates on my work and would love to hear about yours too. 050