Sign in

Onno Eberhard

@onnoeberhard.com
562 followers 298 following 64 posts

PhD Student in Tübingen (MPI-IS & Uni Tü), interested in reinforcement learning. Wir müssen wissen. Wir werden wissen. onnoeberhard.com

PostsRepliesMedia
Onno Eberhard @onnoeberhard.com · 06/07/2026
Our main result is that Committed Q-learning converges to the optimal reactive policy in any "rewire-robust" POMDP, which generalizes both 𝑞⋆-realizability and quasi-Markov environments. The reason is that, by adding commitment, Q-learning automatically learns "entrance values"!
100
Onno Eberhard @onnoeberhard.com · 06/07/2026
We propose the Committed Q-learning algorithm, shown here. The only change to the regular Q-learning algorithm is that a new action is only sampled when the observed feature changes. (Remove the green highlighted if-statement and you get regular Q-learning; the behavior policy 𝜋 is fixed.)
100
Onno Eberhard @onnoeberhard.com · 06/07/2026
We prove that in environments which have clearly defined entrance states, which we call "quasi-Markov" environments, this "naive" value function *always* satisfies this "entrance value" property.
110
Onno Eberhard @onnoeberhard.com · 06/07/2026
Now, the interesting thing here is that the value that we assign to the corridor feature is exactly the value of the corridor entrance state! Thus, the greedy policy here is the same that we would get if we had full state information! But, why does this happen? And how can we adapt it to RL?
100
Onno Eberhard @onnoeberhard.com · 06/07/2026
These are of course not the true dynamics (both environment and policy are deterministic!), but if we simply pretend that the features evolve according to these "pseudo-stochastic" dynamics, we can easily write down a value function!
110
Onno Eberhard @onnoeberhard.com · 06/07/2026
Now consider a naive approach: this environment is non-Markovian (if the underlying state is not observable), but we can pretend that it is actually "stochastic". Under the "go right" policy, we see the transition z=1 → z=1 exactly k-1 out of k times and z=1 → z=2 exactly 1 out of k times.
100
Onno Eberhard @onnoeberhard.com · 06/07/2026
If we change the left-hand reward, the optimal action is to go to the left from x=0, but if we use the minimum VE/BE solutions for the suboptimal "go right" policy, we get these values (k is the corridor length). Notice that a greedy policy with respect to these values will again go right!
100
Onno Eberhard @onnoeberhard.com · 06/07/2026
With different underlying values, it makes sense that Q-learning fails: how can we hope assign a single value to the corridor? A principled starting point would be to use the minimum value error or Bellman error solution. However, these solutions do not guarantee policy improvement!
110
Onno Eberhard @onnoeberhard.com · 06/07/2026
The corridor environment is a very simple POMDP in which all "corridor states" are represented by the same feature/observation (z=1). Q-learning fails here (see plot above), since this environment is not 𝑞⋆-realizable (the different corridor states have different optimal values).
110
Onno Eberhard @onnoeberhard.com · 06/07/2026
I am in Seoul at ICML to present our newest paper "Commit to the Bit: Reactive Reinforcement Learning Done Right". We show that finding reactive policies in POMDPs is easier than previously thought and that the ubiquitous 𝑞⋆-realizability assumption is stronger than necessary. 🧵
2202
Onno Eberhard @onnoeberhard.com · 09/09/2025
A cute little animation: a critically damped harmonic oscillator becomes unstable with integral control if the gain is too high. Here, at K_i = 2, a Hopf bifurcation occurs: two poles of the transfer function enter the right-hand s-plane and the closed-loop system becomes unstable.
032
Onno Eberhard @onnoeberhard.com · 16/07/2025
Memory traces are trivially simple to implement, and we ran some experiments that demonstrate that they are an effective drop-in replacement for sliding windows ("frame stacking") in deep reinforcement learning.
140
Onno Eberhard @onnoeberhard.com · 16/07/2025
However, if we allow larger values of 𝜆, then we do find environments where memory traces are considerably more powerful than sliding windows!
130
Onno Eberhard @onnoeberhard.com · 16/07/2025
Our second result goes the other way: when 𝜆 < 1/2, then there is no environment where memory traces are more efficient than sliding windows. In other words, if 𝜆 < 1/2, then learning with sliding windows and memory traces is equivalent!
110
Onno Eberhard @onnoeberhard.com · 16/07/2025
Using this result, we can finally compare learning with sliding windows to learning with memory traces! Our first result shows that there is no environment where sliding windows are generally more efficient than memory traces (even when restricting to 𝜆 < 1/2).
110
Onno Eberhard @onnoeberhard.com · 16/07/2025
The "resolution" of a function class is given by its Lipschitz constant. We thus consider the function class ℱ = {𝑓 ∘ 𝑧 ∣ 𝑓 : 𝒵 → ℝ, 𝑓 is 𝐿-Lipschitz}. This allows us to bound the metric entropy. (The constant 𝑑_λ is the Minkowski dimension of 𝒵 if 𝜆 < 1/2.)
120
Onno Eberhard @onnoeberhard.com · 16/07/2025
Without forgetting, the learning is intractable: it is equivalent to keeping the complete history. However, to distinguish histories that differ only far in the past, we need to "zoom in" a lot, as shown here.
110
Onno Eberhard @onnoeberhard.com · 16/07/2025
What about memory traces? Here, I am visualizing the space 𝒵 of all possible memory traces for the case where there are only 3 possible (one-hot) observations, 𝒴 = {a, b, c}. We can show that, if 𝜆 < 1/2, then memory traces preserve all information of the complete history! Nothing is forgotten!
110
Onno Eberhard @onnoeberhard.com · 16/07/2025
We focus on the problem of policy evaluation with offline data where the environment ℰ is a hidden Markov model, and we assume that the observation space 𝒴 is one-hot. Thus, given a function class ℱ, our goal is to find the function 𝑓 ∈ ℱ that minimizes the return error.
110
Onno Eberhard @onnoeberhard.com · 16/07/2025
I am in Vancouver at ICML, and tomorrow I will present our newest paper "Partially Observable Reinforcement Learning with Memory Traces". We argue that eligibility traces are more effective than sliding windows as a memory mechanism for RL in POMDPs. 🧵
36012
Onno Eberhard @onnoeberhard.com · 13/06/2025
Great talk by @claireve.bsky.social about our joint work on memory traces this morning. Come join me at poster 94 if you want to know more! #RLDM2025
1112
Onno Eberhard @onnoeberhard.com · 04/06/2025
With this increased sample efficiency, the algorithm can even tackle high-dimensional, non-smooth, and stochastic MuJoCo environments, as shown here.
120
Onno Eberhard @onnoeberhard.com · 04/06/2025
In this algorithm, the Jacobians are estimated independently at every iteration. However, if the learning rate is not too large (and the dynamics are smooth), then we can be more sample efficient by bootstrapping the Jacobian estimates using recursive least squares.
120
Onno Eberhard @onnoeberhard.com · 04/06/2025
This strategy manages to learn highly effective open-loop controllers, like this one that swings up an inverted pendulum.
120
Onno Eberhard @onnoeberhard.com · 04/06/2025
How should we estimate the Jacobians? They determine how the next state changes if a state or action are perturbed. We simply perturb all actions randomly and fit a linear regression model to the transitions.
120
Onno Eberhard @onnoeberhard.com · 04/06/2025
In RL, we don't know the system dynamics, so we cannot evaluate the Jacobians reqiured by Pontryagin's equations. We prove that it is possible to replace them with estimates and still keep convergence guarantees.
130
Onno Eberhard @onnoeberhard.com · 04/06/2025
How can we optimize an open-loop controller? The fundamental tool is *Pontryagin's principle* which tells us how to compute the gradient of the open-loop objective function.
130
Onno Eberhard @onnoeberhard.com · 04/06/2025
In open-loop control, the actions do not depend of the system's state: there is no feedback policy. Instead, a complete sequence of actions is fixed before the interaction begins. This makes it less powerful than closed loop methods, but optimization can be much easier!
130
Onno Eberhard @onnoeberhard.com · 04/06/2025
I'm flying to Michigan today to present our new paper "A Pontryagin Perspective on Reinforcement Learning" at L4DC, where it has been nominated for the Best Paper Award! We ask the question: is it possible to learn an open-loop controller via RL? 🧵
3245