Marcel Hussing @marcelhussing.bsky.social · 03/09/2026🤖 Coming to CoRL 2026: The first workshop on imperfect robotics data! We’ll explore how to collect, curate & use failures 💢, OOD events❓, and unexpected human behavior 🤔 - data every lab produces but rarely documents, shares, or reuses. Check out workshop.oopsie-data.com 142
Marcel Hussing @marcelhussing.bsky.social · 01/08/2026I've tried to be a good reviewer, read all the papers carefully, and formed my own view. I'm getting page long, clearly LLM generated rebuttals. It is not my job to prompt your LLM into convincing me. What are people's solutions to this? Should I message the AC? #NeurIPS2026 010
Reposted by Marcel HussingClaas Voelcker @cvoelcker.bsky.social · 30/07/2026This is a robot failing to grasp a ball. Almost every robot lab produces clips like this daily… and almost all of them get thrown away. This is the most abundant but underused resource in robot learning. We’re collecting all of it now as “OopsieData”, please join us at oopsie-data.com! (1/13) 17413
Marcel Hussing @marcelhussing.bsky.social · 23/07/2026At this point I'm not even sure... maybe I prefer LLM reviews. 040
Marcel Hussing @marcelhussing.bsky.social · 22/07/2026It's this time of yearstatic.klipy.comMr Bean: Rowan Atkinson WaitingALT: Mr Bean: Rowan Atkinson Waiting 010
Marcel Hussing @marcelhussing.bsky.social · 29/06/2026Post is a bit delayed but last week I passed my dissertation defense on algorithm output stability in RL. 🎉 140
Marcel Hussing @marcelhussing.bsky.social · 26/05/2026This was a very fun project. In Behavior-Consistent Deep RL, we provide a method that aligns the behavior of independently trained policies. It turns out, this works even in high dimensional spaces. Here are 6 seeds of Humanoids (all ca same return). (left) Baseline (right) Ours. 092
Marcel Hussing @marcelhussing.bsky.social · 22/05/2026🚨 New Preprint Alert: Behavior-Consistent Deep Reinforcement Learning 🚨 TLDR: We introduce an approach that achieves behavioral similarity across independent algorithm executions in continuous state-action space deep RL. 1294
Marcel Hussing @marcelhussing.bsky.social · 26/04/2026Why do all LLMs predict 27 as their favorite number? There may be a principled explanation. Learn more at Agents in the Wild at #ICLR2026. @ericeaton.bsky.social, me, @surbhigoel.bsky.social, @mkearnsphilly.bsky.social, @aaroth.bsky.social, @sikatasengupta.bsky.social, @optimistsinc.bsky.social 2122
Marcel Hussing @marcelhussing.bsky.social · 25/04/2026Time for round number two. Stop by #4406 to learn about stability guarantees in RL. #ICLR2026 @ericeaton.bsky.social @mkearnsphilly.bsky.social @aaroth.bsky.social @sikatasengupta.bsky.social @optimistsinc.bsky.social 092
Marcel Hussing @marcelhussing.bsky.social · 25/04/2026At #ICLR2026 presenting our first poster in the morning on Relative Entropy Pathwise Policy Optimization. Stop by at #4613. 🧑🎓 @cvoelcker.bsky.social, @axelbrunnbauer.bsky.social, Michal Naumann, Pieter Abbeel, @ericeaton.bsky.social, Radu Grosu, @sologen.bsky.social @igilitschenski.bsky.social 0164
Reposted by Marcel HussingJHU Computer Science @jhucompsci.bsky.social · 21/04/2026In “Replicable Reinforcement Learning with Linear Function Approximation,” @optimistsinc.bsky.social, @marcelhussing.bsky.social, @mkearnsphilly.bsky.social, @aaroth.bsky.social, @sikatasengupta.bsky.social, & more develop replicable methods for linear function approximation in RL: (5/12)arxiv.orgReplicable Reinforcement Learning with Linear Function ApproximationReplication of experimental results has been a challenge faced by many scientific disciplines, including the field of machine learning. Recent work on the theory of machine learning has formalized rep... 173
Marcel Hussing @marcelhussing.bsky.social · 20/04/2026Greetings from Rio, killing some time before #ICLR2026 @iclr-conf.bsky.social 040
Marcel Hussing @marcelhussing.bsky.social · 03/04/2026Why is the default option in the rebuttal acknowledgements at ICML to accept the paper? I'm very confused on how to use these buttons. Can someone explain them to me? 020
Reposted by Marcel HussingAaron Roth @aaroth.bsky.social · 12/03/2026Neural networks are highly non-convex, so approximate error minimizers need not look anything like each other in parameter space. But we show that nevertheless (for many model sizes) approximate error minimizers must closely agree in function/prediction space despite this! 1107
Marcel Hussing @marcelhussing.bsky.social · 10/03/2026It's this time of the year again: your baselines cannot be PPO and SAC. 121
Reposted by Marcel HussingAaron Roth @aaroth.bsky.social · 09/03/2026Michael @mkearnsphilly.bsky.social ) and I wrote a blog post about our experiences using AI for research, and our thoughts on what these developments will mean for research, publication, and education: www.amazon.science/blog/how-ai-...amazon.scienceHow AI is changing the nature of mathematical researchWhat machine learning theorists learned using AI agents to generate proofs — and what comes next. 13013
Marcel Hussing @marcelhussing.bsky.social · 08/03/2026I'm so glad that so many research problems are finally being treated as first class citizens rather than afterthoughts. 🤔 010
Marcel Hussing @marcelhussing.bsky.social · 06/03/2026Too many papers sound like this Hierarchical Context-Aware Diffusion-Transformer Meta-World-Model Reinforcement Learning with Causally Disentangled Preference-Aligned Self-Supervised Compositional Multi-Scale Latent Skill Priors for Long-Horizon Generalist Decision Making 0110
Marcel Hussing @marcelhussing.bsky.social · 05/03/2026One reason I work on replicable and consistent RL is because it is has always been at the top of the list of criteria for reliability. 020
Marcel Hussing @marcelhussing.bsky.social · 03/03/2026Excuse me? Surely telling me that didn't require much thinking. 210
Marcel Hussing @marcelhussing.bsky.social · 03/03/2026I have seen multiple times now that a reviewer said sth like: the proofs are simple -> reject the paper. That is completely counter-productive. A theorem needs to generate new insights. If we learn something new from something simple that should be preferred. Don't believe me? Ask someone famous: 010
Marcel Hussing @marcelhussing.bsky.social · 14/02/2026I’ve been thinking about a practical question and would love some opinions: How do your papers actually get discovered/cited? I was searching for recent work on high update ratio RL and found several very closely related papers tackling the same failure modes we study. None cited our earlier work. 592
Reposted by Marcel HussingIgor Gilitschenski @igilitschenski.bsky.social · 13/02/2026🚀 Excited to share REPPO, a new on-policy RL agent! TL;DR: Replace PPO with REPPO for fewer hyperparameter headaches and more robust training. REPPO, led by @cvoelcker.bsky.social, will be presented at ICLR 2026. How does it work? 🧵👇 12510
Reposted by Marcel HussingLukas Heinrich @lukasheinrich.com · 01/02/2026Scaling Laws in Particle Physics Data! This is a result I've been itching to share and it's finally out. One of the big open questions is how much better AI-based methods at particle colliders can still become. 1/4 3167
Marcel Hussing @marcelhussing.bsky.social · 27/01/2026"Scientific reviewers should have experience publishing scientific work in related areas" is really not that hot of a take. 000
Marcel Hussing @marcelhussing.bsky.social · 26/01/2026Clicking like on any relevant ICLR paper. Encourage people to post their work here more! 040
Marcel Hussing @marcelhussing.bsky.social · 26/01/2026The other paper accepted to @iclr-conf.bsky.social 2026 🇧🇷. Our work on replicable RL sheds some light on how to consistently make decisions in RL. @ericeaton.bsky.social @mkearnsphilly.bsky.social @aaroth.bsky.social @sikatasengupta.bsky.social @optimistsinc.bsky.social 0115
Marcel Hussing @marcelhussing.bsky.social · 26/01/2026Two papers accepted to @iclr-conf.bsky.social 2026! One of the is REPPO, see below! I think it deserves a lot more recognition. Let's chat about it in Rio! 🇧🇷 010
Marcel Hussing @marcelhussing.bsky.social · 26/01/2026Quite disheartening that there isn't a single workshop at ICLR to present my RL work but there several topics that are listed 5 or 6 times just named differently. 020
Marcel Hussing @marcelhussing.bsky.social · 23/01/2026Our number went down by 0.01 but it's very expensive to run so we can't have error bars. Our algorithm is so much better than the rest, new SOTA! 140
Marcel Hussing @marcelhussing.bsky.social · 17/01/2026I can't believe that this paper is not yet used by literally everyone. Claas doing all he can to make your life easier. Check it out. 131
Reposted by Marcel HussingAaron Roth @aaroth.bsky.social · 09/01/2026Excited about a new paper! Multicalibration turns out to be strictly harder than marginal calibration. We prove tight Omega(T^{2/3}) lower bounds for online multicalibration, separating it from online marginal calibration for which better rates were recently discovered. 1225
Marcel Hussing @marcelhussing.bsky.social · 22/12/2025In emergencies, minutes can decide who lives and who dies. Our team is participating in the Triage Challenge, building AI to empower clinicians and prioritize care when resources are thin: prontotriage.com Also recently featured by Meta: ai.meta.com/blog/upenn-d... (1/5) 152
Marcel Hussing @marcelhussing.bsky.social · 22/12/2025Posted about this last week; I feel like it didn't get as much attention as it deserves. We have a new preprint on using diffusion models to generate compositional data. This work was conducted by Quan, a student I advised over the fall. He is currently looking for PhD positions. Check it out! 060
Marcel Hussing @marcelhussing.bsky.social · 19/12/2025We have a new preprint out on iterative generation of compositional robot data. This work was conducted by Quan, one of the students I advised over the semester. Check out the thread! P.S. Quan is currently looking for PhD positions, keep an eye out for him! 041
Marcel Hussing @marcelhussing.bsky.social · 01/12/2025Getting ready to leave for NeurIPS tomorrow morning. ✈️ Let me know if you're around and wanna hang out! 000
Marcel Hussing @marcelhussing.bsky.social · 11/11/2025Me all day todaymedia.tenor.coma man wearing a suit and tie is standing in a field of yellow flowersALT: a man wearing a suit and tie is standing in a field of yellow flowers 030
Marcel Hussing @marcelhussing.bsky.social · 31/10/2025It's mind-boggling to my how many of the papers I reviewed don't cite a single paper that is older than like 2020. It's like we try to collectively forget what people did in the past so we can publish more. 140
Marcel Hussing @marcelhussing.bsky.social · 27/10/2025Am at Columbia today giving a talk about (in part) this work. Have a few hours to kill afterwards. If anyone is around there and wants to chat, DM me. 030
Marcel Hussing @marcelhussing.bsky.social · 26/10/2025I think I posted about it before but never with a thread. We recently put a new preprint on arxiv. 📖 Replicable Reinforcement Learning with Linear Function Approximation 🔗 arxiv.org/abs/2509.08660 In this paper, we study formal replicability in RL with linear function approximation. The... (1/6)arxiv.orgReplicable Reinforcement Learning with Linear Function ApproximationReplication of experimental results has been a challenge faced by many scientific disciplines, including the field of machine learning. Recent work on the theory of machine learning has formalized rep... 2257
Reposted by Marcel HussingClaas Voelcker @cvoelcker.bsky.social · 26/09/2025I have been told I need to get more modern in my paper promotion! github.com/cvoelcker/reppo / arxiv.org/abs/2507.11019 @marcelhussing.bsky.social 1122
Marcel Hussing @marcelhussing.bsky.social · 11/09/2025Super stoked for the New York RL workshop tomorrow. Will be presenting 2 orals: * Replicable Reinforcement Learning with Linear Function Approximation * Relative Entropy Pathwise Policy Optimization We already posted about the 2nd one (below), I'll get to talking about the first one in a bit here. 052
Marcel Hussing @marcelhussing.bsky.social · 16/08/2025(Maybe) unpopular opinion: There should not be *any* new experiments in a rebuttal. A rebuttal is for clarifications and incorrect statements in a review. You should not be allowed to add new content at that point. Either your paper is done or it isn't. It should not be written during rebuttals. 040
Marcel Hussing @marcelhussing.bsky.social · 17/07/2025My PhD journey started with me fine-tuning hparams of PPO which ultimately led to my research on stability. With REPPO, we've made a huge step in the right direction. Stable learning, no tuning on a new benchmark, amazing performance. REPPO has the potential to be the PPO killer we all waited for. 072
Reposted by Marcel HussingClaas Voelcker @cvoelcker.bsky.social · 17/07/2025🔥 Presenting Relative Entropy Pathwise Policy Optimization #REPPO 🔥 Off-policy #RL (eg #TD3) trains by differentiating a critic, while on-policy #RL (eg #PPO) uses Monte-Carlo gradients. But is that necessary? Turns out: No! We show how to get critic gradients on-policy. arxiv.org/abs/2507.11019 2267
Reposted by Marcel HussingClaas Voelcker @cvoelcker.bsky.social · 19/06/2025Works that use #VAML/ #MuZero losses often use deterministic models. But if we want to use stochastic models to measure uncertainty or because we want to leverage current SOTA models such as #transformers and #diffusion, we need to take care! Naively translating the loss functions leads to mistakes! 174