Marcel Hussing @marcelhussing.bsky.social · 07/10/2026Like many, I've thought a lot about the developments in theoretical research. And while I'm not as confident as others here on what to make of them, I think one thing is clear. Recent results are yet another historical marker likely left by RL and PG-style methods. #math #openai 010
Reposted by Marcel HussingClaas Voelcker @cvoelcker.bsky.social · 01/10/2026⏰ More time to submit! ⏰ We hope your drafts are getting into shape and we got good news! The #oopsiedata workshop @ #CoRL2026 🤖 deadline is now October 5th. Polish up your work and send it our way! We are looking forward to all the excellent submissions! 122
Marcel Hussing @marcelhussing.bsky.social · 03/09/2026To find out more about our other efforts surrounding failure and suboptimal datasets, check out the Oopsie-Data project at oopsie-data.comoopsie-data.comHomeTools for collecting, annotating, and managing robotic manipulation failure data — part of the Oopsie Data project. 000
Marcel Hussing @marcelhussing.bsky.social · 03/09/2026This workshop is co-organized with the Oopsie-Data coordination team: @cvoelcker.bsky.social , Siddhant Agarwal, Arpit Bahety, Zhiyuan (Paul) Zhou, @renwang435.bsky.social, Jiahui Chen, Maria Attarian, Carl Qi, @maxbrudolph.bsky.social, Roberto Martín-Martín, Sergey Levine, Peter Stone, Amy Zhang 100
Marcel Hussing @marcelhussing.bsky.social · 03/09/2026🤖 Coming to CoRL 2026: The first workshop on imperfect robotics data! We’ll explore how to collect, curate & use failures 💢, OOD events❓, and unexpected human behavior 🤔 - data every lab produces but rarely documents, shares, or reuses. Check out workshop.oopsie-data.com 142
Marcel Hussing @marcelhussing.bsky.social · 01/08/2026I've tried to be a good reviewer, read all the papers carefully, and formed my own view. I'm getting page long, clearly LLM generated rebuttals. It is not my job to prompt your LLM into convincing me. What are people's solutions to this? Should I message the AC? #NeurIPS2026 010
Reposted by Marcel HussingClaas Voelcker @cvoelcker.bsky.social · 30/07/2026This is a robot failing to grasp a ball. Almost every robot lab produces clips like this daily… and almost all of them get thrown away. This is the most abundant but underused resource in robot learning. We’re collecting all of it now as “OopsieData”, please join us at oopsie-data.com! (1/13) 17413
Marcel Hussing @marcelhussing.bsky.social · 23/07/2026At this point I'm not even sure... maybe I prefer LLM reviews. 040
Marcel Hussing @marcelhussing.bsky.social · 22/07/2026It's this time of yearstatic.klipy.comMr Bean: Rowan Atkinson WaitingALT: Mr Bean: Rowan Atkinson Waiting 010
Marcel Hussing @marcelhussing.bsky.social · 29/06/2026Post is a bit delayed but last week I passed my dissertation defense on algorithm output stability in RL. 🎉 140
Marcel Hussing @marcelhussing.bsky.social · 26/05/2026This was a very fun project. In Behavior-Consistent Deep RL, we provide a method that aligns the behavior of independently trained policies. It turns out, this works even in high dimensional spaces. Here are 6 seeds of Humanoids (all ca same return). (left) Baseline (right) Ours. 092
Marcel Hussing @marcelhussing.bsky.social · 22/05/2026This was a fun collaboration between Penn, Princeton, and UT Austin. 🔗 arxiv.org/abs/2605.21214 🧑🎓 @liv-daliberti.bsky.social @cvoelcker.bsky.social @ben-eysenbach.bsky.social @ericeaton.bsky.socialarxiv.orgBehavior-Consistent Deep Reinforcement LearningReinforcement learning (RL) often exhibits high variance across training runs, leading to unreliable performance and posing a major challenge to deployment in real-world domains. In this work, we addr... 020
Marcel Hussing @marcelhussing.bsky.social · 22/05/2026On standard benchmarks, QED reduces policy divergence across independent runs by up to two orders of magnitude while maintaining competitive performance. 110
Marcel Hussing @marcelhussing.bsky.social · 22/05/2026This motivates Q-value Expectile Disagreement, or QED. QED uses critic disagreement as a proxy for likely behavioral divergence. When disagreement is high, QED increases temperature and keeps the policy exploratory and consistent. When disagreement is low, the policy can specialize. 120
Marcel Hussing @marcelhussing.bsky.social · 22/05/2026But simply increasing entropy is not enough. Too much entropy can hurt optimization and amplify off-policy error, especially when the critic is queried on actions that are poorly supported by the data. 120
Marcel Hussing @marcelhussing.bsky.social · 22/05/2026Our main observation is that maximum-entropy RL already gives us a mechanism for controlling this divergence. Entropy regularization anchors policies toward a shared prior, which can reduce the space of behaviors discovered across independent runs. 120
Marcel Hussing @marcelhussing.bsky.social · 22/05/2026The core challenge is that, during training, two independent runs cannot directly compare their learned policies. Each run only has access to its own data, critics, and policy. So we need a signal within a single run that predicts when executions may diverge. 110
Marcel Hussing @marcelhussing.bsky.social · 22/05/2026In Behavior-Consistent Deep RL, we study consistency in continuous state-action spaces, such as those that arise in robotics. The goal is to make independent training runs produce behaviorally similar policies, not just high-performing ones. 110
Marcel Hussing @marcelhussing.bsky.social · 22/05/2026RL algorithms are notoriously brittle. Even when we use the same algorithm in the same environment, independent runs can produce drastically different behavior. This makes it hard to audit, compare, and reliably update learned policies. 110
Marcel Hussing @marcelhussing.bsky.social · 22/05/2026In a standard ML deployment pipeline, new data is regularly added, a model is retrained, and the updated model is deployed. For this to be useful, updates need to be auditable, reliable, and consistent: we want improvements from new data without unexpected new failure modes. 120
Marcel Hussing @marcelhussing.bsky.social · 22/05/2026🚨 New Preprint Alert: Behavior-Consistent Deep Reinforcement Learning 🚨 TLDR: We introduce an approach that achieves behavioral similarity across independent algorithm executions in continuous state-action space deep RL. 1294
Marcel Hussing @marcelhussing.bsky.social · 27/04/2026I guess I phrased it a bit awkwardly in the hectic of presenting. It's not about the number 27 specifically but rather about the fact that even though these models ostensibly have different architectures and they are highly non convex, they still predict the **same number**. 010
Marcel Hussing @marcelhussing.bsky.social · 27/04/2026In our work we show that for certain model classes (including relu neural networks) loss minimization (not global optimality) implies agreement in prediction space. If you are interested in how we get to that result @aaroth.bsky.social wrote a summary thread here bsky.app/profile/aaro... 140
Marcel Hussing @marcelhussing.bsky.social · 27/04/2026Unclear how this is affected by reasoning and post training but this was a kind of a surprising thing that came up a few years ago. Here is a reddit thread www.reddit.com/r/dataisbeau...reddit.comFrom the dataisbeautiful community on Reddit: LLMs and the number 27: Myth tested with 800 prompts [OC]Explore this post and more from the dataisbeautiful community 120
Marcel Hussing @marcelhussing.bsky.social · 26/04/2026Why do all LLMs predict 27 as their favorite number? There may be a principled explanation. Learn more at Agents in the Wild at #ICLR2026. @ericeaton.bsky.social, me, @surbhigoel.bsky.social, @mkearnsphilly.bsky.social, @aaroth.bsky.social, @sikatasengupta.bsky.social, @optimistsinc.bsky.social 2122
Marcel Hussing @marcelhussing.bsky.social · 25/04/2026Time for round number two. Stop by #4406 to learn about stability guarantees in RL. #ICLR2026 @ericeaton.bsky.social @mkearnsphilly.bsky.social @aaroth.bsky.social @sikatasengupta.bsky.social @optimistsinc.bsky.social 092
Marcel Hussing @marcelhussing.bsky.social · 25/04/2026At #ICLR2026 presenting our first poster in the morning on Relative Entropy Pathwise Policy Optimization. Stop by at #4613. 🧑🎓 @cvoelcker.bsky.social, @axelbrunnbauer.bsky.social, Michal Naumann, Pieter Abbeel, @ericeaton.bsky.social, Radu Grosu, @sologen.bsky.social @igilitschenski.bsky.social 0164
Reposted by Marcel HussingJHU Computer Science @jhucompsci.bsky.social · 21/04/2026In “Replicable Reinforcement Learning with Linear Function Approximation,” @optimistsinc.bsky.social, @marcelhussing.bsky.social, @mkearnsphilly.bsky.social, @aaroth.bsky.social, @sikatasengupta.bsky.social, & more develop replicable methods for linear function approximation in RL: (5/12)arxiv.orgReplicable Reinforcement Learning with Linear Function ApproximationReplication of experimental results has been a challenge faced by many scientific disciplines, including the field of machine learning. Recent work on the theory of machine learning has formalized rep... 173
Marcel Hussing @marcelhussing.bsky.social · 20/04/2026Greetings from Rio, killing some time before #ICLR2026 @iclr-conf.bsky.social 040
Marcel Hussing @marcelhussing.bsky.social · 03/04/2026Why is the default option in the rebuttal acknowledgements at ICML to accept the paper? I'm very confused on how to use these buttons. Can someone explain them to me? 020
Marcel Hussing @marcelhussing.bsky.social · 25/03/2026That is even if the baseline reported 80% on a similar task 030
Reposted by Marcel HussingAaron Roth @aaroth.bsky.social · 12/03/2026Neural networks are highly non-convex, so approximate error minimizers need not look anything like each other in parameter space. But we show that nevertheless (for many model sizes) approximate error minimizers must closely agree in function/prediction space despite this! 1107
Marcel Hussing @marcelhussing.bsky.social · 10/03/2026It's this time of the year again: your baselines cannot be PPO and SAC. 121
Reposted by Marcel HussingAaron Roth @aaroth.bsky.social · 09/03/2026Michael @mkearnsphilly.bsky.social ) and I wrote a blog post about our experiences using AI for research, and our thoughts on what these developments will mean for research, publication, and education: www.amazon.science/blog/how-ai-...amazon.scienceHow AI is changing the nature of mathematical researchWhat machine learning theorists learned using AI agents to generate proofs — and what comes next. 13013
Marcel Hussing @marcelhussing.bsky.social · 08/03/2026I'm so glad that so many research problems are finally being treated as first class citizens rather than afterthoughts. 🤔 010
Marcel Hussing @marcelhussing.bsky.social · 06/03/2026Too many papers sound like this Hierarchical Context-Aware Diffusion-Transformer Meta-World-Model Reinforcement Learning with Causally Disentangled Preference-Aligned Self-Supervised Compositional Multi-Scale Latent Skill Priors for Long-Horizon Generalist Decision Making 0110
Marcel Hussing @marcelhussing.bsky.social · 05/03/2026One reason I work on replicable and consistent RL is because it is has always been at the top of the list of criteria for reliability. 020
Marcel Hussing @marcelhussing.bsky.social · 03/03/2026Excuse me? Surely telling me that didn't require much thinking. 210
Marcel Hussing @marcelhussing.bsky.social · 03/03/2026I have seen multiple times now that a reviewer said sth like: the proofs are simple -> reject the paper. That is completely counter-productive. A theorem needs to generate new insights. If we learn something new from something simple that should be preferred. Don't believe me? Ask someone famous: 010
Marcel Hussing @marcelhussing.bsky.social · 16/02/2026I think it's relatively simple. One side already has the job they want and the other side needs citations to get that job. And everyone tells me advertising work is how you get citations. I have been tempted to go back because on bsky, interactions have become fewer and fewer. Not going to though... 060
Marcel Hussing @marcelhussing.bsky.social · 15/02/2026Yet somehow every now and then a paper becomes very popular even though its findings are similar to those of many others. This paper gets cited while the others don't. Was it just luck? 010
Marcel Hussing @marcelhussing.bsky.social · 15/02/2026While I agree with the sentiment let me play devil's advocate. You only get invited to give talks if your work is already well known. Conferences have become too large to even find relevant people. And social media posts are only marginally relevant if you don't already have a large following. 220
Marcel Hussing @marcelhussing.bsky.social · 14/02/2026I understand that that is okay but at some point it honestly becomes disheartening if you constantly have to reach out to people. There are others who don't seem to have to; what are they doing differently? 100
Marcel Hussing @marcelhussing.bsky.social · 14/02/2026I’m trying to understand whether this is mostly about keyword mismatch, venue visibility, social media, etc. For example, when I search terms like “high update ratio RL” on Scholar, our papers show up near the top. scholar.google.com/scholar?hl=e... Where are things going wrong?scholar.google.com 110
Marcel Hussing @marcelhussing.bsky.social · 14/02/2026I’ve been thinking about a practical question and would love some opinions: How do your papers actually get discovered/cited? I was searching for recent work on high update ratio RL and found several very closely related papers tackling the same failure modes we study. None cited our earlier work. 592
Reposted by Marcel HussingIgor Gilitschenski @igilitschenski.bsky.social · 13/02/2026🚀 Excited to share REPPO, a new on-policy RL agent! TL;DR: Replace PPO with REPPO for fewer hyperparameter headaches and more robust training. REPPO, led by @cvoelcker.bsky.social, will be presented at ICLR 2026. How does it work? 🧵👇 12510
Reposted by Marcel HussingLukas Heinrich @lukasheinrich.com · 01/02/2026Scaling Laws in Particle Physics Data! This is a result I've been itching to share and it's finally out. One of the big open questions is how much better AI-based methods at particle colliders can still become. 1/4 3167