Sign in

Lucas Alegre

@lnalegre.bsky.social
2K followers 171 following 33 posts

Professor at INF - @ufrgs.br | Ph.D. in Computer Science. I am interested in multi-policy reinforcement learning (RL) algorithms. Personal page: lucasalegre.github.io

PostsRepliesMedia
Lucas Alegre @lnalegre.bsky.social · 11/03/2026
Wow, I just realized SUMO-RL reached 1,000 stars on GitHub! 🥳 I created SUMO-RL while I was an undergrad, getting familiar with RL. Traffic signal control is a very cool real-world problem in which RL shines. I'm glad that the community still benefits from it! github.com/LucasAlegre/...
github.com
GitHub - LucasAlegre/sumo-rl: Reinforcement Learning environments for Traffic Signal Control with SUMO. Compatible with Gymnasium, PettingZoo, and popular RL libraries.
Reinforcement Learning environments for Traffic Signal Control with SUMO. Compatible with Gymnasium, PettingZoo, and popular RL libraries. - LucasAlegre/sumo-rl
020
Lucas Alegre @lnalegre.bsky.social · 04/12/2025
It was very fun to present our last paper "Constructing an Optimal Behavior Basis for the Option Keyboard" at NeurIPS this week! Paper: openreview.net/pdf?id=D4gOo... #neurips2025
030
Lucas Alegre @lnalegre.bsky.social · 14/10/2025
Sure, only ~2 weeks to review 5 papers for ICLR. I am sure that all reviewers will have sufficient time to write careful and thoughtful reviews in the following weeks, since they have nothing else to do. It is insane to expect a fair reviewing system in these terms.
020
Lucas Alegre @lnalegre.bsky.social · 03/09/2025
It is really cool to see our work on multi-step GPI being cited in this amazing survey! :) proceedings.neurips.cc/paper_files/...
proceedings.neurips.cc
060
Lucas Alegre @lnalegre.bsky.social · 01/08/2025
On average I have a good score, but it has happened to me before to have 3/4 reviewers accepting the paper, and 1 negative reviewer convincing the AC to reject.
000
Lucas Alegre @lnalegre.bsky.social · 01/08/2025
And now I got the classic rebuttal response: "I have no concerns with the paper, all the theory is great, but since you did not run experiments in expensive domains with image-based environments, I will not increase my score". The goal of experiments is to validate the claims! Not to beat Atari!
120
Lucas Alegre @lnalegre.bsky.social · 24/07/2025
I got the classic NeurIPS reviews "why did you not compare with [completely unrelated method whose comparison would not help support any of the paper's claim]?" Questioning myself whether I should spend my weekend running this useless experiment or if I should argue with the reviewer.
130
Lucas Alegre @lnalegre.bsky.social · 20/06/2025
Finally, reporting only IQM may compromise scientific transparency and fairness, as it can mask poor or unstable performance. Agarwal et al. (2021), who introduced IQM in this context, recommend using it in conjunction with other statistics rather than as a standalone measure.
010
Lucas Alegre @lnalegre.bsky.social · 20/06/2025
Yes, Interquartile Mean (IQM) is a robust statistic that reduces the influence of outliers. But it does not by itself provide a clear and fair analysis of performance. In particular, IQM does not capture the full distribution of returns and may hide important information about variability and risk.
110
Lucas Alegre @lnalegre.bsky.social · 20/06/2025
While I really like the paper "Deep Reinforcement Learning at the Edge of the Statistical Precipice" (openreview.net/forum?id=uqv...), I have seen papers evaluating performance using only the IQM metric and claiming that it is a fairer metric than the mean based on this paper, which is simply wrong.
openreview.net
Deep Reinforcement Learning at the Edge of the Statistical Precipice
Our findings call for a change in how we report performance on benchmarks when using only a few runs, for which we present more reliable protocols accompanied with an open-source library.
161
Lucas Alegre @lnalegre.bsky.social · 02/06/2025
This work was done during my time as an intern at Disney Research Zürich. It was amazing and really fun to develop this idea with the Robotics Team!
020
Lucas Alegre @lnalegre.bsky.social · 02/06/2025
Check out AMOR now on arXiv: Paper: arxiv.org/abs/2505.23708 Full Video: youtube.com/watch?v=gQid... #SIGGRAPH2025 #RL #robotics
120
Lucas Alegre @lnalegre.bsky.social · 02/06/2025
A base policy with uniform weights might fail on challenging motions, but with a few weight tweaks, it nails them. Like this double spin. 🌀😵‍💫 Curious how tuning weights mid-motion can help improve the sim-to-real gap and unlock dynamic, expressive behaviors?
110
Lucas Alegre @lnalegre.bsky.social · 02/06/2025
AMOR trains a single policy conditioned on reward weights and motion context, letting you fine-tune the reward after training. Want smoother motions? Better accuracy? Just adjust the weights — no retraining needed!
110
Lucas Alegre @lnalegre.bsky.social · 02/06/2025
We are excited to share our #SIGGRAPH2025 paper, “AMOR: Adaptive Character Control through Multi-Objective Reinforcement Learning”! Lucas Alegre*, Agon Serifi*, Ruben Grandia, David Müller, Espen Knoop, Moritz Baecher
130
Lucas Alegre @lnalegre.bsky.social · 02/06/2025
Annoyed by having to retrain your entire policy just because your reward weights did not quite work on the real robot? 🤖 www.youtube.com/watch?v=gQid...
youtube.com
AMOR: Adaptive Character Control through Multi-Objective Reinforcement Learning
YouTube video by DisneyResearchHub
161
Lucas Alegre @lnalegre.bsky.social · 30/05/2025
Thank you, Peter! :)
010
Lucas Alegre @lnalegre.bsky.social · 29/05/2025
I'm really glad to have been selected as one of the ICML 2025 Top Reviewers! Too bad I won't be able to go since my last submission was not accepted, even with scores Accept, Accept, Weak Accept, and Weak Reject 🫠
060
Lucas Alegre @lnalegre.bsky.social · 17/03/2025
Last week, I was at @khipu-ai.bsky.social in Santiago, Chile. It was really amazing to see so many great speakers and researchers from Latin America together!
041
Reposted by Lucas Alegre
Pablo Samuel Castro @pcastr.bsky.social · 05/03/2025
RL is so back! (well, for some of us, it never really left) awards.acm.org/about/2024-t...
awards.acm.org
Andrew Barto and Richard Sutton are the recipients of the 2024 ACM A.M. Turing Award for developing the conceptual and algorithmic foundations of reinforcement learning.
Andrew Barto and Richard Sutton as the recipients of the 2024 ACM A.M. Turing Award for developing the conceptual and algorithmic foundations of reinforcement learning. In a series of papers beginning...
17212
Lucas Alegre @lnalegre.bsky.social · 16/02/2025
Thank you! link.springer.com/article/10.1... This paper is a great start point!
link.springer.com
A practical guide to multi-objective reinforcement learning and planning - Autonomous Agents and Multi-Agent Systems
Real-world sequential decision-making tasks are generally complex, requiring trade-offs between multiple, often conflicting, objectives. Despite this, the majority of research in reinforcement learnin...
020
Lucas Alegre @lnalegre.bsky.social · 16/02/2025
Thank you! 😊
010
Lucas Alegre @lnalegre.bsky.social · 16/02/2025
Finally, I would like to thank my advisors, Prof. Ana Bazzan and Prof. Bruno C. da Silva; Prof. Ann Nowé who received me at VUB for my PhD stay; and Disney Research Zürich, where I interned. I am very grateful to everyone with that I had the chance to collaborate in all such amazing projects! 💙
010
Lucas Alegre @lnalegre.bsky.social · 16/02/2025
I believe all these contributions open room for many interesting ideas for multi-policy RL methods. Especially in transfer learning (SFs&GPI) and multi-objective RL settings! 🚀
110
Lucas Alegre @lnalegre.bsky.social · 16/02/2025
* MO-Gymnasium (github.com/Farama-Found...) is a library of MORL environments; and * MORL Baselines (github.com/LucasAlegre/...) is a library of MORL algorithms. Both have become standards in MORL research and have over 100k downloads in the past year!
github.com
GitHub - Farama-Foundation/MO-Gymnasium: Multi-objective Gymnasium environments for reinforcement learning
Multi-objective Gymnasium environments for reinforcement learning - Farama-Foundation/MO-Gymnasium
110
Lucas Alegre @lnalegre.bsky.social · 16/02/2025
Besides the theoretical and algorithmic contributions, we also introduced an open-source toolkit for MORL research! NeurIPS D&B 2023 Paper - openreview.net/pdf?id=jfwRL...
openreview.net
110
Lucas Alegre @lnalegre.bsky.social · 16/02/2025
Next, we further explored how to leverage approximate models of the environment to improve zero-shot policy transfer. Our method, ℎ-GPI, interpolates between model-free GPI and fully model-based planning as a function of the planning horizon ℎ. NeurIPS 2023 Paper - openreview.net/pdf?id=KFj0Q...
openreview.net
110
Lucas Alegre @lnalegre.bsky.social · 16/02/2025
Next, we further explored these ideas and introduced two novel MORL algorithms that exploit GPI to increase sample efficiency in MORL: GPI-LS and GPI-PD. AAMAS'23 paper: tinyurl.com/aamas23
tinyurl.com
Sample-Efficient Multi-Objective Learning via Generalized Policy Improvement Prioritization | Proceedings of the 2023 International Conference on Autonomous Agents and Multiagent Systems
You will be notified whenever a record that you have chosen has been cited.
110
Lucas Alegre @lnalegre.bsky.social · 16/02/2025
By exploiting these connections, we introduced SFOLS, a method that is capable of constructing a set of policies and combining them via GPI with the guarantee of obtaining the optimal policy for any novel linearly-expressible tasks! ICML'22 paper: proceedings.mlr.press/v162/alegre2...
proceedings.mlr.press
Optimistic Linear Support and Successor Features as a Basis for Optimal Policy Transfer
In many real-world applications, reinforcement learning (RL) agents might have to solve multiple tasks, each one typically modeled via a reward function. If reward functions are expressed linearly,...
110
Lucas Alegre @lnalegre.bsky.social · 16/02/2025
It all started when we discovered and introduced connections between Successor Features and multi-objective RL (MORL):
proceedings.mlr.press
Optimistic Linear Support and Successor Features as a Basis for Optimal Policy Transfer
In many real-world applications, reinforcement learning (RL) agents might have to solve multiple tasks, each one typically modeled via a reward function. If reward functions are expressed linearly,...
110
Lucas Alegre @lnalegre.bsky.social · 16/02/2025
I thought it would be a good idea to make a thread highlighting the main contributions of my Ph.D! 🧵
110
Lucas Alegre @lnalegre.bsky.social · 16/02/2025
I am happy to announce that I successfully defended my PhD, entitled “Sample-Efficieny Multi-Task and Multi-Objective Reinforcement Learning by Combining Multiple Behaviors”! 🎉 These last years have been extremely fun, and I am very lucky to have collaborated with and met so many great people😄
3161
Reposted by Lucas Alegre
Ed Clark @clarkai.bsky.social · 22/11/2024
Another must read for reinforcement learning. Answers many key questions for researchers; -Do I need multiple training runs? -How do I report model confidence? -And a great section on common mistakes to fend off reviewer 2 🧪 #DRL #reinforcementlearning #AI arxiv.org/abs/2304.01315
arxiv.org
Empirical Design in Reinforcement Learning
Empirical design in reinforcement learning is no small task. Running good experiments requires attention to detail and at times significant computational resources. While compute resources available p...
1385
Lucas Alegre @lnalegre.bsky.social · 24/11/2024
Anyone else dislike the idea of papers being almost completely rewritten from scratch during ICLR rebuttals? This period should be used to address minor issues. I find a bit unfair authors expecting the reviewers to increase score when half of the relevant results were only shown during rebuttal.
150
Lucas Alegre @lnalegre.bsky.social · 16/09/2024
I would be happy to be added too :)
210