Reposted by Onno Eberhard Michal Valko @misovalko.bsky.social · 08/08/2026Explaining a poster with both hands, which is how you know the result matters. 171
Onno Eberhard @onnoeberhard.com · 07/07/2026Come see our poster tomorrow morning (#4307 in Hall A)! Paper, code, and more can be found at onnoeberhard.com/q-commit. Joint work with @claireve.bsky.social and Michael Muhelebach.onnoeberhard.com Commit to the Bit: Reactive Reinforcement Learning Done Right · Onno Eberhard ML & Mathematics 031
Onno Eberhard @onnoeberhard.com · 06/07/2026Thanks! Nice observation, I’m actually working on a follow-up project right now that tries to do exactly this 😄 120
Onno Eberhard @onnoeberhard.com · 06/07/2026This is joint work with @claireve.bsky.social and Michael Muehlebach. If you are at ICML, please come to our poster on Wednesday morning (poster #4307 in Hall A)! Paper, code, and more can be found at onnoeberhard.com/q-commit.onnoeberhard.com Commit to the Bit: Reactive Reinforcement Learning Done Right · Onno Eberhard ML & Mathematics 030
Onno Eberhard @onnoeberhard.com · 06/07/2026Our main result is that Committed Q-learning converges to the optimal reactive policy in any "rewire-robust" POMDP, which generalizes both 𝑞⋆-realizability and quasi-Markov environments. The reason is that, by adding commitment, Q-learning automatically learns "entrance values"! 100
Onno Eberhard @onnoeberhard.com · 06/07/2026We propose the Committed Q-learning algorithm, shown here. The only change to the regular Q-learning algorithm is that a new action is only sampled when the observed feature changes. (Remove the green highlighted if-statement and you get regular Q-learning; the behavior policy 𝜋 is fixed.) 100
Onno Eberhard @onnoeberhard.com · 06/07/2026We prove that in environments which have clearly defined entrance states, which we call "quasi-Markov" environments, this "naive" value function *always* satisfies this "entrance value" property. 110
Onno Eberhard @onnoeberhard.com · 06/07/2026Now, the interesting thing here is that the value that we assign to the corridor feature is exactly the value of the corridor entrance state! Thus, the greedy policy here is the same that we would get if we had full state information! But, why does this happen? And how can we adapt it to RL? 100
Onno Eberhard @onnoeberhard.com · 06/07/2026These are of course not the true dynamics (both environment and policy are deterministic!), but if we simply pretend that the features evolve according to these "pseudo-stochastic" dynamics, we can easily write down a value function! 110
Onno Eberhard @onnoeberhard.com · 06/07/2026Now consider a naive approach: this environment is non-Markovian (if the underlying state is not observable), but we can pretend that it is actually "stochastic". Under the "go right" policy, we see the transition z=1 → z=1 exactly k-1 out of k times and z=1 → z=2 exactly 1 out of k times. 100
Onno Eberhard @onnoeberhard.com · 06/07/2026If we change the left-hand reward, the optimal action is to go to the left from x=0, but if we use the minimum VE/BE solutions for the suboptimal "go right" policy, we get these values (k is the corridor length). Notice that a greedy policy with respect to these values will again go right! 100
Onno Eberhard @onnoeberhard.com · 06/07/2026With different underlying values, it makes sense that Q-learning fails: how can we hope assign a single value to the corridor? A principled starting point would be to use the minimum value error or Bellman error solution. However, these solutions do not guarantee policy improvement! 110
Onno Eberhard @onnoeberhard.com · 06/07/2026The corridor environment is a very simple POMDP in which all "corridor states" are represented by the same feature/observation (z=1). Q-learning fails here (see plot above), since this environment is not 𝑞⋆-realizable (the different corridor states have different optimal values). 110
Onno Eberhard @onnoeberhard.com · 06/07/2026I am in Seoul at ICML to present our newest paper "Commit to the Bit: Reactive Reinforcement Learning Done Right". We show that finding reactive policies in POMDPs is easier than previously thought and that the ubiquitous 𝑞⋆-realizability assumption is stronger than necessary. 🧵 2202
Reposted by Onno EberhardNeha Binish @nbinish.bsky.social · 04/05/2026Excited that our paper is finally out in @natneuro.nature.com 🎉 Huge thanks to all my co-authors, especially @jonasterlau.bsky.social & @randolph-helfrich.bsky.social for the support and great teamwork 🍀 1254
Reposted by Onno EberhardRandolph Helfrich @randolph-helfrich.bsky.social · 04/05/2026Very excited to share our latest paper led by @nbinish.bsky.social in @natneuro.nature.com We demonstrate how a communication subspace channels higher-dimensional PFC dynamics into lower-D motor activity to enable efficient behavior using human iEEG. www.nature.com/articles/s41... #neuroskyencenature.comA communication subspace relays context-dependent actions from human prefrontal to motor cortex - Nature NeuroscienceContext-dependent behavior selects actions according to task demands. Using direct brain recordings in humans, Binish et al. uncover how coordinated population activity efficiently channels informatio... 311044
Reposted by Onno EberhardSIGBOVIK @harryqbovik.bsky.social · 03/03/2026The deadline for all SIGBOVIK 2026 papers has officially been extended to March 18! Enjoy the extra procrastination time, and maybe consider starting to write your papers! 1134
Reposted by Onno EberhardEWRL @ewrl-org.bsky.social · 25/11/2025Exciting workshop for RL enthusiasts in Mannheim! 👇 Workshop on Reinforcement Learning 2026, taking place on 𝐅𝐞𝐛𝐫𝐮𝐚𝐫𝐲 𝟔, 𝟐𝟎𝟐𝟔, at the 𝐔𝐧𝐢𝐯𝐞𝐫𝐬𝐢𝐭𝐲 𝐨𝐟 𝐌𝐚𝐧𝐧𝐡𝐞𝐢𝐦, Germany. Participation in the workshop is 𝐟𝐫𝐞𝐞 𝐨𝐟 𝐜𝐡𝐚𝐫𝐠𝐞! Check the program and register: www.wim.uni-mannheim.de/doering/conf... 283
Reposted by Onno EberhardClaire Vernade @claireve.bsky.social · 16/10/2025Nicolo Cesa-Bianchi and Matteo Papini are putting together a great unconference workshop at the @ellis.eu day at @euripsconf.bsky.social If you want to talk about RL, causality, bandits, online learning, join us there on December 2nd sites.google.com/view/ilir-wo...sites.google.comILIR Workshop, Dec 2, 2025This workshop covers current research topics in reinforcement learning and causality, and in particular questions at the interface of these research areas. Of particular interest this year are also qu... 1207
Reposted by Onno EberhardMichela Petriconi @michelapetriconi.bsky.social · 22/09/2025I had such a great time helping organize EWRL 2025 with an amazing team 🎉 Loved being part of it and meeting so many passionate reinforcement learning enthusiasts! @ewrl18.bsky.social 081
Reposted by Onno EberhardMax-Planck-Gesellschaft @maxplanck.de · 19/09/2025Truly chuffed for our fearless food physicists @mpipks.bsky.social + collabs from AT @istaresearch.bsky.social, IT & ES who won this year’s Ig Nobel - the #NobelPrize of hearts❤️for cracking the science of perfect pasta !🍝Kudos to all for intrepidly consuming lots of cheese in the name of science!😋the-scientist.comThe Secret to a Smooth Pasta Sauce Wins Ig Nobel PrizeItalian researchers studied how the ingredients of the traditional Roman dish cacio e pepe emulsify into a creamy sauce, winning the 2025 Physics Ig Nobel Prize. 06417
Onno Eberhard @onnoeberhard.com · 12/09/2025I wrote a short post on our newest ICML paper addressed at people who are not experts in machine learning. Check it out! 093
Onno Eberhard @onnoeberhard.com · 09/09/2025A cute little animation: a critically damped harmonic oscillator becomes unstable with integral control if the gain is too high. Here, at K_i = 2, a Hopf bifurcation occurs: two poles of the transfer function enter the right-hand s-plane and the closed-loop system becomes unstable. 032
Reposted by Onno EberhardEWRL @ewrl-org.bsky.social · 13/08/2025📣Registration for EWRL is now open📣 Register now 👇 and join us in Tübingen for 3 days (17th-19th September) full of inspiring talks, posters and many social activities to push the boundaries of the RL community!site.pheedloop.comPheedLoopPheedLoop: Hybrid, In-Person & Virtual Event Software 084
Reposted by Onno EberhardGeorg Martius @gmartius.bsky.social · 16/07/2025I am going to present the poster during the next poster session. 11am Wed. Poster W #707 052
Reposted by Onno EberhardEugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 16/07/2025I really, really like this paper and as an open question, would love to see it tested on more memory benchmarks 1312
Reposted by Onno EberhardClaire Vernade @claireve.bsky.social · 16/07/2025Onno and I will be presenting our poster at # W1005 tomorrow (Wed) morning. He made a great thread about it, come chat with us about POMDP theory :) 0195
Onno Eberhard @onnoeberhard.com · 16/07/2025This is joint work with @claireve.bsky.social and Michael Muehlebach. If you are at ICML, please come to our poster tomorrow morning (W-1005, Tuesday, 11am-1:30pm). Paper, code, and more can be found at onnoeberhard.com/memory-traces.onnoeberhard.com Partially Observable Reinforcement Learning with Memory Traces · Onno Eberhard ML & Mathematics 160
Onno Eberhard @onnoeberhard.com · 16/07/2025Memory traces are trivially simple to implement, and we ran some experiments that demonstrate that they are an effective drop-in replacement for sliding windows ("frame stacking") in deep reinforcement learning. 140
Onno Eberhard @onnoeberhard.com · 16/07/2025However, if we allow larger values of 𝜆, then we do find environments where memory traces are considerably more powerful than sliding windows! 130
Onno Eberhard @onnoeberhard.com · 16/07/2025Our second result goes the other way: when 𝜆 < 1/2, then there is no environment where memory traces are more efficient than sliding windows. In other words, if 𝜆 < 1/2, then learning with sliding windows and memory traces is equivalent! 110
Onno Eberhard @onnoeberhard.com · 16/07/2025Using this result, we can finally compare learning with sliding windows to learning with memory traces! Our first result shows that there is no environment where sliding windows are generally more efficient than memory traces (even when restricting to 𝜆 < 1/2). 110
Onno Eberhard @onnoeberhard.com · 16/07/2025The "resolution" of a function class is given by its Lipschitz constant. We thus consider the function class ℱ = {𝑓 ∘ 𝑧 ∣ 𝑓 : 𝒵 → ℝ, 𝑓 is 𝐿-Lipschitz}. This allows us to bound the metric entropy. (The constant 𝑑_λ is the Minkowski dimension of 𝒵 if 𝜆 < 1/2.) 120
Onno Eberhard @onnoeberhard.com · 16/07/2025Without forgetting, the learning is intractable: it is equivalent to keeping the complete history. However, to distinguish histories that differ only far in the past, we need to "zoom in" a lot, as shown here. 110
Onno Eberhard @onnoeberhard.com · 16/07/2025What about memory traces? Here, I am visualizing the space 𝒵 of all possible memory traces for the case where there are only 3 possible (one-hot) observations, 𝒴 = {a, b, c}. We can show that, if 𝜆 < 1/2, then memory traces preserve all information of the complete history! Nothing is forgotten! 110
Onno Eberhard @onnoeberhard.com · 16/07/2025We are interested in efficiently learning an accurate value estimate. Statistical learning theory tells us that efficient learning is easier if the *metric entropy* 𝐻(ℱ) is small. For window memory, the function class ℱ is ℱₘ ≐ {𝑓 ∘ winₘ ∣ 𝑓: 𝒴ᵐ → ℝ}, and the metric entropy is 𝐻(ℱₘ) ∈ Θ(|𝒴|ᵐ). 120
Onno Eberhard @onnoeberhard.com · 16/07/2025We focus on the problem of policy evaluation with offline data where the environment ℰ is a hidden Markov model, and we assume that the observation space 𝒴 is one-hot. Thus, given a function class ℱ, our goal is to find the function 𝑓 ∈ ℱ that minimizes the return error. 110
Onno Eberhard @onnoeberhard.com · 16/07/2025While most theoretical work on memory in RL focuses on sliding windows of observations, winₘ(𝑦ₜ, 𝑦ₜ₋₁, … ) ≐ (𝑦ₜ, 𝑦ₜ₋₁, …, 𝑦ₜ₋ₘ₊₁), we analyze the effectiveness of *memory traces*, exponential moving averages of observations: 𝑧(𝑦ₜ, 𝑦ₜ₋₁, … ) = 𝜆𝑧(𝑦ₜ₋₁, 𝑦ₜ₋₂, … ) + (1 − 𝜆)𝑦ₜ. 120
Onno Eberhard @onnoeberhard.com · 16/07/2025I am in Vancouver at ICML, and tomorrow I will present our newest paper "Partially Observable Reinforcement Learning with Memory Traces". We argue that eligibility traces are more effective than sliding windows as a memory mechanism for RL in POMDPs. 🧵 36012
Onno Eberhard @onnoeberhard.com · 13/06/2025This result should thus also transfer to approximate memory traces. However, the connection between memory traces and truncated histories only applies if the forgetting factor lambda is less than 1/2. The case of lambda > 1/2 is more interesting, but the connection to AIS is much less clear to me. 120
Onno Eberhard @onnoeberhard.com · 13/06/2025I believe that this case is indeed closely related to AIS. Our analysis describes a close connection between approximate memory traces and truncated histories. Under some conditions (e.g. gamma-observability), truncated histories constitute approximate information states (if I understand correctly). 110
Onno Eberhard @onnoeberhard.com · 13/06/2025I am not sure if there is a way to relate the case where these conditions are not met to AIS. However, we study the behavior of Lipschitz continuous functions of memory traces, which is closely related to quantizing the space of memory traces. 110
Onno Eberhard @onnoeberhard.com · 13/06/2025Interesting question! In the paper, we identify very general conditions under which the memory trace is an exact information state. For example, if the set of observations is linearly independent, then it suffices for the forgetting factor lambda to be rational. 110
Onno Eberhard @onnoeberhard.com · 13/06/2025For those not at RLDM, the paper (and the poster) can be found at onnoeberhard.com/memory-traces. 📄onnoeberhard.com Partially Observable Reinforcement Learning with Memory Traces · Onno Eberhard ML & Mathematics 030
Onno Eberhard @onnoeberhard.com · 13/06/2025Great talk by @claireve.bsky.social about our joint work on memory traces this morning. Come join me at poster 94 if you want to know more! #RLDM2025 1112
Reposted by Onno EberhardClaire Vernade @claireve.bsky.social · 13/06/2025This is a joint work with @onnoeberhard.com and Michael Mühlebach, and the poster will be presented by Onno tonight at #RLDM, and later in July at @icmlconf.bsky.social 231
Reposted by Onno EberhardClaire Vernade @claireve.bsky.social · 13/06/2025This morning at #RLDM I talked about Memory Traces, a simple representation for POMDPs that’s probably good at remembering old observations. arxiv.org/abs/2503.15200 2243