Sign in

Axel Brunnbauer

@axelbrunnbauer.bsky.social
50 followers 158 following 6 posts

RL @ Amazon RIVR

PostsRepliesMedia
Reposted by Axel Brunnbauer
Marcel Hussing @marcelhussing.bsky.social · 26/05/2026
This was a very fun project. In Behavior-Consistent Deep RL, we provide a method that aligns the behavior of independently trained policies. It turns out, this works even in high dimensional spaces. Here are 6 seeds of Humanoids (all ca same return). (left) Baseline (right) Ours.
092
Reposted by Axel Brunnbauer
Marcel Hussing @marcelhussing.bsky.social · 25/04/2026
At #ICLR2026 presenting our first poster in the morning on Relative Entropy Pathwise Policy Optimization. Stop by at #4613. 🧑‍🎓 @cvoelcker.bsky.social, @axelbrunnbauer.bsky.social, Michal Naumann, Pieter Abbeel, @ericeaton.bsky.social, Radu Grosu, @sologen.bsky.social @igilitschenski.bsky.social
0164
Reposted by Axel Brunnbauer
Igor Gilitschenski @igilitschenski.bsky.social · 13/02/2026
🚀 Excited to share REPPO, a new on-policy RL agent! TL;DR: Replace PPO with REPPO for fewer hyperparameter headaches and more robust training. REPPO, led by @cvoelcker.bsky.social, will be presented at ICLR 2026. How does it work? 🧵👇
12510
Reposted by Axel Brunnbauer
Claas Voelcker @cvoelcker.bsky.social · 26/01/2026
Or… you can chat with us in 🇧🇷 Rio 🇧🇷 as we are going to @iclr-conf.bsky.social to present our paper!!!
191
Reposted by Axel Brunnbauer
Claas Voelcker @cvoelcker.bsky.social · 17/01/2026
🤔 Want to use REPPO (cvoelcker.de/projects/rep...) but hate jax? 🤔 😮 Want to have stable on-policy RL without filling your GPU with an enormous replay buffer? 😮 🤖 Are you a roboticist and just want your RL code to run? 🤖 🎉 Fear not, we started adding new REPPO versions! 🎉 github.com/cvoelcker/rs...
cvoelcker.de
Relative Entropy Pathwise Policy Optimization | Claas A. Voelcker
A simple, whitespace theme for academics. Based on [*folio](https://github.com/bogoli/-folio) design.
2164
Axel Brunnbauer @axelbrunnbauer.bsky.social · 03/10/2025
This blog post is a nice complementary, behind-the-scenes extra on our recent work about on-policy pathwise gradient algorithms. @cvoelcker.bsky.social went the extra mile, and wrote this piece to provide some more context on the design decisions behind REPPO!
130
Reposted by Axel Brunnbauer
Claas Voelcker @cvoelcker.bsky.social · 02/10/2025
cvoelcker.de/blog/2025/re... I finally gave in and made a nice blog post about my most recent paper. This was a surprising amount of work, so please be nice and go read it!
media.tenor.com
a close up of a sad cat with the words pleeeaasse written below it
ALT: a close up of a sad cat with the words pleeeaasse written below it
0297
Reposted by Axel Brunnbauer
Claas Voelcker @cvoelcker.bsky.social · 16/09/2025
Big if true 🤫: #REPPO works on Atari as well 😱 👾 🚀 Some tuning is still needed, but we are seeing results roughly on par with #PQN. If you want to test out #REPPO (atari is not integrated due to issues with envpool and jax version), check out github.com/cvoelcker/re... #reinforcementlearning
A lonely return curve on the ALE game Qbert-v5 for the REPPO algorithm
171
Reposted by Axel Brunnbauer
Marcel Hussing @marcelhussing.bsky.social · 11/09/2025
Super stoked for the New York RL workshop tomorrow. Will be presenting 2 orals: * Replicable Reinforcement Learning with Linear Function Approximation * Relative Entropy Pathwise Policy Optimization We already posted about the 2nd one (below), I'll get to talking about the first one in a bit here.
052
Reposted by Axel Brunnbauer
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 18/07/2025
I’ve been hearing about this paper from Claas for a while now, the fact that they aren’t tuning per benchmark is a killer sign. Also, check out the wall clock plots!
1201
Reposted by Axel Brunnbauer
Marcel Hussing @marcelhussing.bsky.social · 17/07/2025
My PhD journey started with me fine-tuning hparams of PPO which ultimately led to my research on stability. With REPPO, we've made a huge step in the right direction. Stable learning, no tuning on a new benchmark, amazing performance. REPPO has the potential to be the PPO killer we all waited for.
072
Reposted by Axel Brunnbauer
Claas Voelcker @cvoelcker.bsky.social · 17/07/2025
🔥 Presenting Relative Entropy Pathwise Policy Optimization #REPPO 🔥 Off-policy #RL (eg #TD3) trains by differentiating a critic, while on-policy #RL (eg #PPO) uses Monte-Carlo gradients. But is that necessary? Turns out: No! We show how to get critic gradients on-policy. arxiv.org/abs/2507.11019
GIF showing two plots that symbolize the REPPO algorithm. On the left side, four curves track the return of an optimization function, and on the right side, the optimization paths over the objective function are visualized. The GIF shows that monte-carlo gradient estimators have a high variance and fail to converge, while surrogate function estimators converge smoothly, but might find suboptimal solutions if the surrogate function is imprecise.
2267
Axel Brunnbauer @axelbrunnbauer.bsky.social · 11/02/2025
Our paper on unsupervised environment design for autonomous-driving scenarios was accepted at ICRA! We built a curriculum generator for CARLA which adapts the scenario distribution to the current capabilities of the agent. arxiv.org/abs/2403.17805
arxiv.org
Scenario-Based Curriculum Generation for Multi-Agent Autonomous Driving
The automated generation of diverse and complex training scenarios has been an important ingredient in many complex learning tasks. Especially in real-world application domains, such as autonomous dri...
000
Axel Brunnbauer @axelbrunnbauer.bsky.social · 20/12/2024
Excited to announce that our paper "Scalable Offline Reinforcement Learning for Mean Field Games" has been accepted at #AAMAS2025! 🚀 We propose Off-MMD, an offline RL algorithm for learning equilibrium policies in MFGs from static datasets. arxiv.org/abs/2410.17898
arxiv.org
Scalable Offline Reinforcement Learning for Mean Field Games
Reinforcement learning algorithms for mean-field games offer a scalable framework for optimizing policies in large populations of interacting agents. Existing methods often depend on online interactio...
010