Sign in

Gaspard Lambrechts

@gsprd.be
2.2K followers 657 following 35 posts

Postdoctoral researcher working on RL in POMDP at McGill and Mila - gsprd.be

PostsRepliesMedia
Reposted by Gaspard Lambrechts
ELLIS @ellis.eu · 04/03/2026
📣 Reinforcement Learning Summer School is returning to Milan in 2026! Co-organized with @ellisunitmilan.bsky.social & designed for Master's and PhD students on RL theory, multi-agent systems, RL & LLMs, real-world applications... 📍 Milan 🇮🇹 📅 3-12 June ⏰ Apply by 27 March 🔗 bit.ly/4b2Plhp
02411
Gaspard Lambrechts @gsprd.be · 20/02/2026
Congratulations to the hardworking folks at UPenn! Thank you Edward for including me and for all the nice discussions. 🌐 penn-pal-lab.github.io/aawr 📝 openreview.net/forum?id=Rkd... 💻 github.com/penn-pal-lab... More theory details in Appendix A-E and on slide 30 (orbi.uliege.be/handle/2268/...)
010
Gaspard Lambrechts @gsprd.be · 20/02/2026
As seen from the results and videos, AAWR improves significantly on (i) foundation policies, (ii) behavior cloning policies, and (iii) AWR policies, providing good policies even in partially observable environments with non Markovian inputs.
110
Gaspard Lambrechts @gsprd.be · 20/02/2026
It is the case here, where we use additional cameras, position estimates, or bounding boxes from pretrained models. These features i are used as additional input of the critic Q(i, z, a) to provide a better advantage estimate and policy improvement direction.
100
Gaspard Lambrechts @gsprd.be · 20/02/2026
In fact, it is a common assumption in asymmetric RL, which distinguishes the execution information from the training information. In practice, while we do not always know the exact state, it is common to have more information available about the state at training time.
120
Gaspard Lambrechts @gsprd.be · 20/02/2026
Moreover, (s, z) is shown to be the Markovian state of an equivalent MDP, allowing us to rely on the Bellman equation and TD learning instead of MC learning. Now, how realistic is it to assume that we know the state s in addition to the input z?
100
Gaspard Lambrechts @gsprd.be · 20/02/2026
Unfortunately, we show that we cannot just learn the symmetric critic Q(z, a) = E[G | z, a]. Instead, we need an asymmetric critic Q(s, z, a) = E[G | s, z, a] for a valid policy iteration. This is because, unlike for policy gradients, the AWR objective is not linear in Q.
100
Gaspard Lambrechts @gsprd.be · 20/02/2026
To learn a good policy (for this specific input z), we may want to rely on existing RL algorithms such as policy gradient or policy iteration. Here, because we perform offline-to-online training, we rely on AWR, a policy iteration algorithm going offline to online seamlessly.
100
Gaspard Lambrechts @gsprd.be · 20/02/2026
When learning in real-world scenarios, it is common to have constraints on the input available to the policy at execution time (e.g., last observation only, wrist camera only, etc). In the general case (POMDP), the input z is a function of the observation history h: z = f(h).
100
Gaspard Lambrechts @gsprd.be · 20/02/2026
At NeurIPS, we presented asymmetric RL algo: Asymmetric Advantage Weighted Regression (AAWR). This time, the goal is to learn a policy pi(a | z) whose input z = f(h) is not necessarily Markovian. It is useful in robotics, for example. So, how to adapt RL in POMDP to non Markovian input? 🧵
youtu.be
AAWR: Real World RL of Active Perception Behaviors @ NeurIPS2025
YouTube video by Jie Wang
120
Gaspard Lambrechts @gsprd.be · 06/10/2025
Interestingly, this distribution can be learned from off-policy samples with a TD-like update. And this, even when encouraging the visitation of features of future states (possibly aliased). arxiv.org/abs/2412.06655
arxiv.org
Off-Policy Maximum Entropy RL with Future State and Action Visitation Measures
Maximum entropy reinforcement learning integrates exploration into policy learning by providing additional intrinsic rewards proportional to the entropy of some distribution. In this paper, we propose...
010
Gaspard Lambrechts @gsprd.be · 06/10/2025
4) Off-Policy Maximum Entropy RL with Future State and Action Visitation Measures. With Adrien Bolland and Damien Ernst, we propose a new intrinsic reward. Instead of encouraging visiting states uniformly, we encourage visiting *future* states uniformly, from every state.
120
Gaspard Lambrechts @gsprd.be · 06/10/2025
This view offers interesting insights for the design of intrinsic rewards, by providing four criteria. arxiv.org/abs/2402.00162
arxiv.org
Behind the Myth of Exploration in Policy Gradients
In order to compute near-optimal policies with policy-gradient algorithms, it is common in practice to include intrinsic exploration terms in the learning objective. Although the effectiveness of thes...
110
Gaspard Lambrechts @gsprd.be · 06/10/2025
3) Behind the Myth of Exploration in Policy Gradients. With Adrien Bolland and Damien Ernst, we decided to frame the exploration problem for policy-gradient methods from the optimization point of view.
110
Gaspard Lambrechts @gsprd.be · 06/10/2025
By adapting a finite-time bound, we uncover an interesting tradeoff between informativeness of the additional information and complexity of the resulting value function. openreview.net/forum?id=wNV...
openreview.net
Informed Asymmetric Actor-Critic: Theoretical Insights and Open...
Reinforcement learning in partially observable environments requires agents to make decisions under uncertainty, based on incomplete and noisy observations. Asymmetric actor-critic methods improve...
140
Gaspard Lambrechts @gsprd.be · 06/10/2025
2) Informed Asymmetric Actor-Critic: Theoretical Insights and Open Questions. With Daniel Ebi and Damien Ernst, we looked for a reason why asymmetric actor-critic was performing better, even when using RNN-based policies with the full observation history as input (no aliasing).
110
Gaspard Lambrechts @gsprd.be · 06/10/2025
In AsymAC, while the policy maintains an agent state based on observations only, the critic also takes the state as input. Its better performance is linked to eventual "aliasing" in the agent state, hurting TD learning in the symmetric case only. arxiv.org/abs/2501.19116
arxiv.org
A Theoretical Justification for Asymmetric Actor-Critic Algorithms
In reinforcement learning for partially observable environments, many successful algorithms have been developed within the asymmetric learning paradigm. This paradigm leverages additional state inform...
120
Gaspard Lambrechts @gsprd.be · 06/10/2025
1) A Theoretical Justification for Asymmetric Actor-Critic Algorithms. With Damien Ernst and Aditya Mahajan, we looked for a reason why asymmetric actor-critic algorithms are performing better than their symmetric counterparts.
110
Gaspard Lambrechts @gsprd.be · 06/10/2025
At #EWRL, we presented 4 papers, which we summarize below. - A Theoretical Justification for AsymAC Algorithms. - Informed AsymAC: Theoretical Insights and Open Questions. - Behind the Myth of Exploration in Policy Gradients. - Off-Policy MaxEntRL with Future State-Action Visitation Measures.
141
Reposted by Gaspard Lambrechts
Théo Vincent @theo-vincent.bsky.social · 19/07/2025
Had an amazing time presenting my research @cohereforai.bsky.social yesterday 🎤 In case you could not attend, feel free to check it out 👉 youtu.be/RCA22JWiiY8?...
youtu.be
Théo Vincent - Optimizing the Learning Trajectory of Reinforcement Learning Agents
YouTube video by Cohere
073
Reposted by Gaspard Lambrechts
Claire Vernade @claireve.bsky.social · 17/07/2025
Such an inspiring talk by @arkrause.bsky.social at #ICML today. The role of efficient exploration in Scientific discovery is fundamental and I really like how Andreas connects the dots with RL (theory).
0152
Gaspard Lambrechts @gsprd.be · 16/07/2025
At #ICML2025, we will present a theoretical justification for the benefits of « asymmetric actor-critic » algorithms (#W1008 Wednesday at 11am). 📝 Paper: hdl.handle.net/2268/326874 💻 Blog: damien-ernst.be/2025/06/10/a...
ICML poster of the paper « A Theoretical Justification for Asymmetric Actor-Critic Algorithms » by Gaspard Lambrechts, Damien Ernst and Aditya Mahajan.
083
Reposted by Gaspard Lambrechts
Riccardo Zamboni @ricczamboni.bsky.social · 08/07/2025
🌟🌟Good news for the explorers🗺️! Next week we will present our paper “Enhancing Diversity in Parallel Agents: A Maximum Exploration Story” with V. De Paola, @mircomutti.bsky.social and M. Restelli at @icmlconf.bsky.social! (1/N)
141
Gaspard Lambrechts @gsprd.be · 11/07/2025
Last week, I gave an invited talk on "asymmetric reinforcement learning" at the BeNeRL workshop. I was happy to draw attention to this niche topic, which I think can be useful to any reinforcement learning researcher. Slides: hdl.handle.net/2268/333931.
063
Gaspard Lambrechts @gsprd.be · 13/06/2025
Two months after my PhD defense on RL in POMDP, I finally uploaded the final version of my thesis :) You can find it here: hdl.handle.net/2268/328700 (manuscript and slides). Many thanks to my advisors and to the jury members.
Cover page of the PhD thesis "Reinforcement Learning in Partially Observable Markov Decision Processes: Learning to Remember the Past by Learning to Predict the Future" by Gaspard Lambrechts
082
Gaspard Lambrechts @gsprd.be · 09/06/2025
TL;DR: Do not make the problem harder than it is! Using state information during training is provably better. 📝 Paper: arxiv.org/abs/2501.19116 🎤 Talk: orbi.uliege.be/handle/2268/... A warm thank to Aditya Mahajan for welcoming me at McGill University and for his precious supervision.
arxiv.org
A Theoretical Justification for Asymmetric Actor-Critic Algorithms
In reinforcement learning for partially observable environments, many successful algorithms have been developed within the asymmetric learning paradigm. This paradigm leverages additional state inform...
041
Gaspard Lambrechts @gsprd.be · 09/06/2025
While this work has considered fixed feature z = f(h) with linear approximators, we discuss possible generalizations in the conclusion. Despite not matching the usual recurrent actor-critic setting, this analysis still provides insights into the effectiveness of asymmetric actor-critic algorithms.
100
Gaspard Lambrechts @gsprd.be · 09/06/2025
The conclusion is that asymmetric learning is less sensitive to aliasing than symmetric learning. Now, what is aliasing exactly? The aliasing and inference terms arise from z = f(h) not being Markovian. They can be bounded by the difference between the approximate p(s|z) and exact p(s|h) beliefs.
100
Gaspard Lambrechts @gsprd.be · 09/06/2025
Now, as far as the actor suboptimality is concerned, we obtained the following finite-time bounds. In addition to the average critic error, which is also present in the actor bound, the symmetric actor-critic algorithm suffers from an additional "inference term".
Theorem showing the finite-time suboptimality bound for the asymmetric and symmetric actor-critic algorithms. The asymmetric algorithm has four terms: the natural actor-critic term, the gradient estimation term, the residual gradient term, and the average critic error. The symmetric algorithm has an additional term: the inference term.
100
Gaspard Lambrechts @gsprd.be · 09/06/2025
By adapting the finite-time bound from the symmetric setting to the asymmetric setting, we obtain the following error bounds for the critic estimates. The symmetric temporal difference learning algorithm has an additional "aliasing term".
Theorem showing the finite-time error bound for the asymmetric and symmetric temporal difference learning algorithms. The asymmetric algorithm has three terms: the temporal difference learning term, the function approximation term, and the bootstrapping shift term. The symmetric algorithm has an additional term: the aliasing term.
100
Gaspard Lambrechts @gsprd.be · 09/06/2025
While this algorithm is valid/unbiased (Baisero & Amato, 2022), a theoretical justification for its benefit is still missing. Does it really learn faster than symmetric learning? In this paper, we provide theoretical evidence for this, based on an adapted finite-time analysis (Cayci et al., 2024).
Title page of the paper "A Theoretical Justification for Asymmetric Actor-Critic Algorithms", written by Gaspard Lambrechts, Damien Ernst and Aditya Mahajan.
100
Gaspard Lambrechts @gsprd.be · 09/06/2025
However, with actor-critic algorithms, it can be noticed that the critic is not needed at execution! As a result, the state can be an input of the critic, which becomes Q(s, z, a) in the asymmetric setting instead of Q(z, a) in the symmetric setting.
Figure showing the policy being passed the feature z = f(h) of the history, and the critic being passed both the feature z = f(h) of the history and the state s as input. The asymmetric critic and the policy are used together to form the sample policy-gradient expression.
100
Gaspard Lambrechts @gsprd.be · 09/06/2025
In a POMDP, the goal is to find an optimal policy π(a|z) that maps a feature z = f(h) of the history h to an action a. In a privileged POMDP, the state can be used to learn a policy π(a|z) faster. But note that the state cannot be an input of the policy, since it is not available at execution.
Figure showing the history h being compressed into a feature z = f(h) for then being passed to a policy g(a | z).
100
Gaspard Lambrechts @gsprd.be · 09/06/2025
Typically, classical RL methods assume: - MDP: full state observability (too optimistic), - POMDP: partial state observability (too pessimistic). Instead, asymmetric RL methods assume: - Privileged POMDP: asymmetric state observability (full at training, partial at execution).
Table showing that the MDP assumes full state observability both during training and execution and that the POMDP assumes partial state observability both during training and execution, while the privileged POMDP assumes full state observability during training but partial state observability during execution.
100
Gaspard Lambrechts @gsprd.be · 09/06/2025
📝 Our paper "A Theoretical Justification for Asymmetric Actor-Critic Algorithms" was accepted at #ICML! Never heard of "asymmetric actor-critic" algorithms? Yet, many successful #RL applications use them (see image). But these algorithms are not fully understood. Below, we provide some insights.
Slide showing three recent successes of reinforcement learning that have used an asymmetric actor-critic algorithm:
 - Magnetic Control of Tokamak Plasma through Deep RL (Degrave et al., 2022).
 - Champion-Level Drone Racing using Deep RL (Kaufmann et al., 2023).
 - A Super-Human Vision-Based RL Agent in Gran Turismo (Vasco et al., 2024).
1177
Reposted by Gaspard Lambrechts
EWRL @ewrl-org.bsky.social · 26/05/2025
📢 Deadline extended! Submit your work to EWRL — now accepting papers until June 3rd AoE. This year, we're also offering a fast track for papers accepted at other conferences ⚡ Check the website for all the details: euro-workshop-on-reinforcement-learning.github.io/ewrl18/
086
Gaspard Lambrechts @gsprd.be · 03/04/2025
Slydst, my Typst package for making simple slides, just got its 100th star on Github. While I would not advise using Typst for papers yet, its markdown-like syntax allows to create slides in a few minutes, while supporting everything we love from LaTeX: equations. github.com/glambrechts/...
Typst interface showing an example of Slydst code and the resulting slides.
030
Reposted by Gaspard Lambrechts
Gilles Louppe @glouppe.bsky.social · 30/12/2024
📣 Hiring! I am looking for PhD/postdoc candidates to work on foundation models for science at @ULiege, with a special focus on weather and climate systems. 🌏 Three positions are open around deep learning, physics-informed FMs and inverse problems with FMs.
47934
Reposted by Gaspard Lambrechts
Adrien Bolland @adrienbolland.bsky.social · 13/12/2024
Check our work on max entropy RL! We introduce an off-policy method to maximize the entropy of the future state-action visitation distribution, leading to policies that explore effectively and achieve high performance 🎯 Link 📑 arxiv.org/abs/2412.06655 #RL #MaxEntRL #Exploration
arxiv.org
Off-Policy Maximum Entropy RL with Future State and Action Visitation Measures
We introduce a new maximum entropy reinforcement learning framework based on the distribution of states and actions visited by a policy. More precisely, an intrinsic reward function is added to the re...
0114
Gaspard Lambrechts @gsprd.be · 03/12/2024
How come I didn't know about this BeNeRL seminar series? It focuses on practical RL and seems really great! www.benerl.org/seminar-seri... I would have loved to hear Benjamin Eysenbach, Chris Lu and Edward Hu... Next one is on December 19th.
092
Gaspard Lambrechts @gsprd.be · 27/11/2024
Hi Erfan, for now I am studying the theory of asymmetric actor-critic algorithms for POMDP. Afterwards, I plan on working with SSM-based WM, and maybe also on an asymmetric WM for POMDP. What about you?
000
Gaspard Lambrechts @gsprd.be · 26/11/2024
Hi Bluesky! I'm a PhD student researching #RL in #POMDP. If you're also interested in: - sequence models, - world models, - representation learning, - asymmetric learning, - generalization, I'd love to connect, chat, or check out your work! Feel free to reply or message me.
1212