Sign in

Mirco Mutti

@mircomutti.bsky.social
1.5K followers 323 following 36 posts

Reinforcement learning, but without rewards. Postdoc at the Technion. PhD from Politecnico di Milano. muttimirco.github.io

PostsRepliesMedia
Reposted by Mirco Mutti
Clément Canonne @ccanonne.github.io · 01/08/2026
"We're all worried," as what it means to do research (in my field, Theoretical CS) seems to be shifting, and shifting fast. What to do? Senior researchers must lead by example, knowing that not everything will pan out. What I'm suggesting below may not work everywhere, but here's my own advice: 1/
620351
Reposted by Mirco Mutti
Riccardo Zamboni @ricczamboni.bsky.social · 02/04/2026
🚨🚨🚨 Later today I am going to present at @rl-agents-rg.bsky.social’s reading group one research line with @mircomutti.bsky.social that I’m really excited about: behavior compression via unsupervised RL!
062
Reposted by Mirco Mutti
ELLIS @ellis.eu · 04/03/2026
📣 Reinforcement Learning Summer School is returning to Milan in 2026! Co-organized with @ellisunitmilan.bsky.social & designed for Master's and PhD students on RL theory, multi-agent systems, RL & LLMs, real-world applications... 📍 Milan 🇮🇹 📅 3-12 June ⏰ Apply by 27 March 🔗 bit.ly/4b2Plhp
02411
Mirco Mutti @mircomutti.bsky.social · 28/11/2025
Absolutely, come to the poster! Some say Riccardo's aura will be hovering around
010
Reposted by Mirco Mutti
Transactions on Machine Learning Research @tmlrorg.bsky.social · 14/10/2025
As Transactions on Machine Learning Research (TMLR) grows in number of submissions, we are looking for more reviewers and action editors. Please sign up! Only one paper to review at a time and <= 6 per year, reviewers report greater satisfaction than reviewing for conferences!
11112
Reposted by Mirco Mutti
EWRL @ewrl-org.bsky.social · 13/08/2025
📣Registration for EWRL is now open📣 Register now 👇 and join us in Tübingen for 3 days (17th-19th September) full of inspiring talks, posters and many social activities to push the boundaries of the RL community!
site.pheedloop.com
PheedLoop
PheedLoop: Hybrid, In-Person & Virtual Event Software
084
Mirco Mutti @mircomutti.bsky.social · 24/07/2025
Walking around posters at @icmlconf.bsky.social, I was happy to see some buzz around convex RL—a topic I’ve worked on and strongly believe in. Thought I’d share a few ICML papers on this direction. Let’s dive in👇 But first… what is convex RL? 🧵 1/n
151
Mirco Mutti @mircomutti.bsky.social · 15/07/2025
Would you trust a bandit algorithm to make decisions on your health or investments? Common exploration mechanisms are efficient but scary. In our latest work at @icmlconf.bsky.social, we reimagine bandit algorithms to get *efficient* and *interpretable* exploration. A 🧵 below 1/n
130
Mirco Mutti @mircomutti.bsky.social · 09/07/2025
Here we have an original take on how to make the best of parallel data collection for RL. Don't miss the poster at ICML, we're curious to hear what y'all think! Kudos to the awesome students Vincenzo and @ricczamboni.bsky.social for their work under the wise supervision of Marcello.
010
Reposted by Mirco Mutti
Amir-massoud Farahmand @sologen.bsky.social · 09/07/2025
What do we talk about when we talk about the Bellman Optimality Equation? If we think carefully, we are (implicitly) making three claims. #FoundationsOfReinforcementLearning #sneakpeek
First, we claim that there exists a unique value function $\Vopt$ that satisfies the following equation: For any $x \in \XX$, we have
\begin{align*}
	\Vopt(x) =
	\max_{a \in \AA} \left \{ r(x,a) + \gamma \int \PKernel(\dx' | x, a) \Vopt(x') \right \}.
\end{align*}
This claim alone, however, does not show that this $\Vopt$ is the same as $V^\piopt$.

The second claim is that $\Vopt$ is indeed the same as $V^{\piopt}$, the optimal value function when $\pi$ is restricted to be within the space of stationary policies.
This claim alone, however, does not preclude the possibility that we can find an ever more performant policy by going beyond the space of stationary policies.

The third claim is that for discounted continuing MDPs, we can always find a stationary policy that is optimal within the space of all stationary and non-stationary policies.

These three claims together show that the Bellman optimality equation reveals the recursive structure of the optimal value function $\Vopt = V^{\piopt}$. There is no policy, stationary or non-stationary, with a value function better than $\Vopt$, for the class of discounted continuing MDPs.
061
Reposted by Mirco Mutti
Gautam Kamath @gautamkamath.com · 07/07/2025
System is so broken: - researchers write papers no one reads - reviewers don't have time to review, shamed to coauthors, use LLMs instead of reading - authors try to fool said LLMs with prompt injection - evaling researchers based on # of papers (no time to read) Dystopic.
1010610
Reposted by Mirco Mutti
EWRL @ewrl-org.bsky.social · 08/04/2025
Mark your calendars, EWRL is coming to Tübingen! 📅 When? September 17-19, 2025. More news to come soon, stay tuned!
03714
Reposted by Mirco Mutti
Tim van Erven @timvanerven.nl · 08/04/2025
Just enjoyed @mircomutti.bsky.social's seminar talk about interpretable meta-learning of contextual bandit types. The recording is available in case you missed it: youtu.be/pNos7AHGMXw
youtu.be
Theory of Interpretable AI Seminar: Mirco Mutti
YouTube video by Theory of Interpretable AI Seminar
081
Mirco Mutti @mircomutti.bsky.social · 08/04/2025
Happening today! Join us if you want to hear about our take on interpretable exploration for multi-armed bandits. If interested but cannot join, here's the arxiv arxiv.org/abs/2504.04505 Joint work with Jeongyeol, Shie, and @aviv-tamar.bsky.social
001
Reposted by Mirco Mutti
Tim van Erven @timvanerven.nl · 27/03/2025
⏰⏰Theory of Interpretable AI Seminar ⏰⏰ In two weeks, April 8, Mirco Mutti will talk about "A Classification View on Meta Learning Bandits"
063
Mirco Mutti @mircomutti.bsky.social · 17/03/2025
The right review form is: - Summary - Comment - Evaluation Curious of alternative arguments, as it looks like conferences are going in a different direction
020
Mirco Mutti @mircomutti.bsky.social · 20/02/2025
Awesome! Have a look at this thread to see some nice multi-object manipulation results
020
Reposted by Mirco Mutti
Aldo Pacchiano @aldopacchiano.bsky.social · 30/01/2025
[4/5] “A Theoretical Framework for Partially-Observed Reward States in RLHF” develops and analyzes a model for RLHF where we posit the human feedback to be generated by a stateful labeler. @mircomutti.bsky.social
141
Mirco Mutti @mircomutti.bsky.social · 13/12/2024
If interested on our take on addressing inverse RL in large state spaces, go to meet @filippo_lazzati and @alberto_metelli in the poster session 5 #NeurIPS2024 today (paper -> arxiv.org/abs/2406.03812)
152
Reposted by Mirco Mutti
Andrea Celli @acelli.bsky.social · 28/11/2024
I will soon be opening a call for a postdoctoral position in online learning and algorithmic game theory, starting in 2025, funded by my ERC at Bocconi University. If you're interested, feel free to reach out. If you're not personally interested but know someone who might be, please let them know!
184
Mirco Mutti @mircomutti.bsky.social · 25/11/2024
Highly recommended!
020
Reposted by Mirco Mutti
Dr. Angelica Lim @petitegeek.bsky.social · 23/11/2024
This is nice brain candy for the affective computing crowd
011
Reposted by Mirco Mutti
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 09/11/2024
If you're an RL researcher or RL adjacent, pipe up to make sure I've added you here! go.bsky.app/3WPHcHg
527127
Mirco Mutti @mircomutti.bsky.social · 20/11/2024
Hello there! I'm new here and interested in AI -especially reinforcement learning- and keeping up with the latest in research. I'll occasionally share updates on my work and would love to hear about yours too.
050