Sign in

Marcel Hussing

@marcelhussing.bsky.social
3K followers 336 following 214 posts

PhD student at the University of Pennsylvania. Prev, intern at MSR, and Meta FAIR. Interested in reliable and replicable reinforcement learning, robotics and knowledge discovery: marcelhussing.github.io All posts are my own.

PostsRepliesMedia
Marcel Hussing @marcelhussing.bsky.social · 03/09/2026
🤖 Coming to CoRL 2026: The first workshop on imperfect robotics data! We’ll explore how to collect, curate & use failures 💢, OOD events❓, and unexpected human behavior 🤔 - data every lab produces but rarely documents, shares, or reuses. Check out workshop.oopsie-data.com
142
Marcel Hussing @marcelhussing.bsky.social · 01/08/2026
I've tried to be a good reviewer, read all the papers carefully, and formed my own view. I'm getting page long, clearly LLM generated rebuttals. It is not my job to prompt your LLM into convincing me. What are people's solutions to this? Should I message the AC? #NeurIPS2026
010
Reposted by Marcel Hussing
Claas Voelcker @cvoelcker.bsky.social · 30/07/2026
This is a robot failing to grasp a ball. Almost every robot lab produces clips like this daily… and almost all of them get thrown away. This is the most abundant but underused resource in robot learning. We’re collecting all of it now as “OopsieData”, please join us at oopsie-data.com! (1/13)
17413
Marcel Hussing @marcelhussing.bsky.social · 23/07/2026
At this point I'm not even sure... maybe I prefer LLM reviews.
040
Marcel Hussing @marcelhussing.bsky.social · 22/07/2026
It's this time of year
static.klipy.com
Mr Bean: Rowan Atkinson Waiting
ALT: Mr Bean: Rowan Atkinson Waiting
010
Marcel Hussing @marcelhussing.bsky.social · 29/06/2026
Post is a bit delayed but last week I passed my dissertation defense on algorithm output stability in RL. 🎉
140
Marcel Hussing @marcelhussing.bsky.social · 26/05/2026
This was a very fun project. In Behavior-Consistent Deep RL, we provide a method that aligns the behavior of independently trained policies. It turns out, this works even in high dimensional spaces. Here are 6 seeds of Humanoids (all ca same return). (left) Baseline (right) Ours.
092
Marcel Hussing @marcelhussing.bsky.social · 22/05/2026
🚨 New Preprint Alert: Behavior-Consistent Deep Reinforcement Learning 🚨 TLDR: We introduce an approach that achieves behavioral similarity across independent algorithm executions in continuous state-action space deep RL.
1294
Marcel Hussing @marcelhussing.bsky.social · 26/04/2026
Why do all LLMs predict 27 as their favorite number? There may be a principled explanation. Learn more at Agents in the Wild at #ICLR2026. @ericeaton.bsky.social, me, @surbhigoel.bsky.social, @mkearnsphilly.bsky.social, @aaroth.bsky.social, @sikatasengupta.bsky.social, @optimistsinc.bsky.social
2122
Marcel Hussing @marcelhussing.bsky.social · 25/04/2026
Time for round number two. Stop by #4406 to learn about stability guarantees in RL. #ICLR2026 @ericeaton.bsky.social @mkearnsphilly.bsky.social @aaroth.bsky.social @sikatasengupta.bsky.social @optimistsinc.bsky.social
092
Marcel Hussing @marcelhussing.bsky.social · 25/04/2026
At #ICLR2026 presenting our first poster in the morning on Relative Entropy Pathwise Policy Optimization. Stop by at #4613. 🧑‍🎓 @cvoelcker.bsky.social, @axelbrunnbauer.bsky.social, Michal Naumann, Pieter Abbeel, @ericeaton.bsky.social, Radu Grosu, @sologen.bsky.social @igilitschenski.bsky.social
0164
Reposted by Marcel Hussing
JHU Computer Science @jhucompsci.bsky.social · 21/04/2026
In “Replicable Reinforcement Learning with Linear Function Approximation,” @optimistsinc.bsky.social, @marcelhussing.bsky.social, @mkearnsphilly.bsky.social, @aaroth.bsky.social, @sikatasengupta.bsky.social, & more develop replicable methods for linear function approximation in RL: (5/12)
arxiv.org
Replicable Reinforcement Learning with Linear Function Approximation
Replication of experimental results has been a challenge faced by many scientific disciplines, including the field of machine learning. Recent work on the theory of machine learning has formalized rep...
173
Marcel Hussing @marcelhussing.bsky.social · 20/04/2026
Greetings from Rio, killing some time before #ICLR2026 @iclr-conf.bsky.social
040
Marcel Hussing @marcelhussing.bsky.social · 17/04/2026
Getting ready to fly to Rio. 😎🇧🇷
000
Marcel Hussing @marcelhussing.bsky.social · 03/04/2026
Why is the default option in the rebuttal acknowledgements at ICML to accept the paper? I'm very confused on how to use these buttons. Can someone explain them to me?
020
Reposted by Marcel Hussing
Aaron Roth @aaroth.bsky.social · 12/03/2026
Neural networks are highly non-convex, so approximate error minimizers need not look anything like each other in parameter space. But we show that nevertheless (for many model sizes) approximate error minimizers must closely agree in function/prediction space despite this!
1107
Marcel Hussing @marcelhussing.bsky.social · 10/03/2026
It's this time of the year again: your baselines cannot be PPO and SAC.
121
Reposted by Marcel Hussing
Aaron Roth @aaroth.bsky.social · 09/03/2026
Michael @mkearnsphilly.bsky.social ) and I wrote a blog post about our experiences using AI for research, and our thoughts on what these developments will mean for research, publication, and education: www.amazon.science/blog/how-ai-...
amazon.science
How AI is changing the nature of mathematical research
What machine learning theorists learned using AI agents to generate proofs — and what comes next.
13013
Marcel Hussing @marcelhussing.bsky.social · 08/03/2026
I'm so glad that so many research problems are finally being treated as first class citizens rather than afterthoughts. 🤔
010
Marcel Hussing @marcelhussing.bsky.social · 06/03/2026
Too many papers sound like this Hierarchical Context-Aware Diffusion-Transformer Meta-World-Model Reinforcement Learning with Causally Disentangled Preference-Aligned Self-Supervised Compositional Multi-Scale Latent Skill Priors for Long-Horizon Generalist Decision Making
0110
Marcel Hussing @marcelhussing.bsky.social · 05/03/2026
One reason I work on replicable and consistent RL is because it is has always been at the top of the list of criteria for reliability.
020
Marcel Hussing @marcelhussing.bsky.social · 03/03/2026
Excuse me? Surely telling me that didn't require much thinking.
210
Marcel Hussing @marcelhussing.bsky.social · 03/03/2026
I have seen multiple times now that a reviewer said sth like: the proofs are simple -> reject the paper. That is completely counter-productive. A theorem needs to generate new insights. If we learn something new from something simple that should be preferred. Don't believe me? Ask someone famous:
010
Marcel Hussing @marcelhussing.bsky.social · 14/02/2026
I’ve been thinking about a practical question and would love some opinions: How do your papers actually get discovered/cited? I was searching for recent work on high update ratio RL and found several very closely related papers tackling the same failure modes we study. None cited our earlier work.
592
Reposted by Marcel Hussing
Igor Gilitschenski @igilitschenski.bsky.social · 13/02/2026
🚀 Excited to share REPPO, a new on-policy RL agent! TL;DR: Replace PPO with REPPO for fewer hyperparameter headaches and more robust training. REPPO, led by @cvoelcker.bsky.social, will be presented at ICLR 2026. How does it work? 🧵👇
12510
Reposted by Marcel Hussing
Lukas Heinrich @lukasheinrich.com · 01/02/2026
Scaling Laws in Particle Physics Data! This is a result I've been itching to share and it's finally out. One of the big open questions is how much better AI-based methods at particle colliders can still become. 1/4
3167
Marcel Hussing @marcelhussing.bsky.social · 27/01/2026
"Scientific reviewers should have experience publishing scientific work in related areas" is really not that hot of a take.
000
Marcel Hussing @marcelhussing.bsky.social · 26/01/2026
Clicking like on any relevant ICLR paper. Encourage people to post their work here more!
040
Marcel Hussing @marcelhussing.bsky.social · 26/01/2026
The other paper accepted to @iclr-conf.bsky.social 2026 🇧🇷. Our work on replicable RL sheds some light on how to consistently make decisions in RL. @ericeaton.bsky.social @mkearnsphilly.bsky.social @aaroth.bsky.social @sikatasengupta.bsky.social @optimistsinc.bsky.social
0115
Marcel Hussing @marcelhussing.bsky.social · 26/01/2026
Two papers accepted to @iclr-conf.bsky.social 2026! One of the is REPPO, see below! I think it deserves a lot more recognition. Let's chat about it in Rio! 🇧🇷
010
Marcel Hussing @marcelhussing.bsky.social · 26/01/2026
Quite disheartening that there isn't a single workshop at ICLR to present my RL work but there several topics that are listed 5 or 6 times just named differently.
020
Marcel Hussing @marcelhussing.bsky.social · 23/01/2026
Our number went down by 0.01 but it's very expensive to run so we can't have error bars. Our algorithm is so much better than the rest, new SOTA!
140
Marcel Hussing @marcelhussing.bsky.social · 17/01/2026
I can't believe that this paper is not yet used by literally everyone. Claas doing all he can to make your life easier. Check it out.
131
Marcel Hussing @marcelhussing.bsky.social · 13/01/2026
Bringing this back up.
030
Reposted by Marcel Hussing
Aaron Roth @aaroth.bsky.social · 09/01/2026
Excited about a new paper! Multicalibration turns out to be strictly harder than marginal calibration. We prove tight Omega(T^{2/3}) lower bounds for online multicalibration, separating it from online marginal calibration for which better rates were recently discovered.
1225
Marcel Hussing @marcelhussing.bsky.social · 22/12/2025
In emergencies, minutes can decide who lives and who dies. Our team is participating in the Triage Challenge, building AI to empower clinicians and prioritize care when resources are thin: prontotriage.com Also recently featured by Meta: ai.meta.com/blog/upenn-d... (1/5)
152
Marcel Hussing @marcelhussing.bsky.social · 22/12/2025
Posted about this last week; I feel like it didn't get as much attention as it deserves. We have a new preprint on using diffusion models to generate compositional data. This work was conducted by Quan, a student I advised over the fall. He is currently looking for PhD positions. Check it out!
060
Marcel Hussing @marcelhussing.bsky.social · 19/12/2025
We have a new preprint out on iterative generation of compositional robot data. This work was conducted by Quan, one of the students I advised over the semester. Check out the thread! P.S. Quan is currently looking for PhD positions, keep an eye out for him!
041
Marcel Hussing @marcelhussing.bsky.social · 01/12/2025
Getting ready to leave for NeurIPS tomorrow morning. ✈️ Let me know if you're around and wanna hang out!
000
Marcel Hussing @marcelhussing.bsky.social · 11/11/2025
Me all day today
media.tenor.com
a man wearing a suit and tie is standing in a field of yellow flowers
ALT: a man wearing a suit and tie is standing in a field of yellow flowers
030
Marcel Hussing @marcelhussing.bsky.social · 31/10/2025
It's mind-boggling to my how many of the papers I reviewed don't cite a single paper that is older than like 2020. It's like we try to collectively forget what people did in the past so we can publish more.
140
Marcel Hussing @marcelhussing.bsky.social · 27/10/2025
Am at Columbia today giving a talk about (in part) this work. Have a few hours to kill afterwards. If anyone is around there and wants to chat, DM me.
030
Marcel Hussing @marcelhussing.bsky.social · 26/10/2025
I think I posted about it before but never with a thread. We recently put a new preprint on arxiv. 📖 Replicable Reinforcement Learning with Linear Function Approximation 🔗 arxiv.org/abs/2509.08660 In this paper, we study formal replicability in RL with linear function approximation. The... (1/6)
arxiv.org
Replicable Reinforcement Learning with Linear Function Approximation
Replication of experimental results has been a challenge faced by many scientific disciplines, including the field of machine learning. Recent work on the theory of machine learning has formalized rep...
2257
Reposted by Marcel Hussing
Claas Voelcker @cvoelcker.bsky.social · 26/09/2025
I have been told I need to get more modern in my paper promotion! github.com/cvoelcker/reppo / arxiv.org/abs/2507.11019 @marcelhussing.bsky.social
Happy guy sad guy meme with sad text: USE PPO AND TUNE HYPERPARAMETER FOR WEEKS and happy text: USE REPPO AND GET A POLICY
1122
Marcel Hussing @marcelhussing.bsky.social · 11/09/2025
Super stoked for the New York RL workshop tomorrow. Will be presenting 2 orals: * Replicable Reinforcement Learning with Linear Function Approximation * Relative Entropy Pathwise Policy Optimization We already posted about the 2nd one (below), I'll get to talking about the first one in a bit here.
052
Marcel Hussing @marcelhussing.bsky.social · 16/08/2025
(Maybe) unpopular opinion: There should not be *any* new experiments in a rebuttal. A rebuttal is for clarifications and incorrect statements in a review. You should not be allowed to add new content at that point. Either your paper is done or it isn't. It should not be written during rebuttals.
040
Marcel Hussing @marcelhussing.bsky.social · 08/08/2025
New ChatGPT data just dropped
0386
Marcel Hussing @marcelhussing.bsky.social · 17/07/2025
My PhD journey started with me fine-tuning hparams of PPO which ultimately led to my research on stability. With REPPO, we've made a huge step in the right direction. Stable learning, no tuning on a new benchmark, amazing performance. REPPO has the potential to be the PPO killer we all waited for.
072
Reposted by Marcel Hussing
Claas Voelcker @cvoelcker.bsky.social · 17/07/2025
🔥 Presenting Relative Entropy Pathwise Policy Optimization #REPPO 🔥 Off-policy #RL (eg #TD3) trains by differentiating a critic, while on-policy #RL (eg #PPO) uses Monte-Carlo gradients. But is that necessary? Turns out: No! We show how to get critic gradients on-policy. arxiv.org/abs/2507.11019
GIF showing two plots that symbolize the REPPO algorithm. On the left side, four curves track the return of an optimization function, and on the right side, the optimization paths over the objective function are visualized. The GIF shows that monte-carlo gradient estimators have a high variance and fail to converge, while surrogate function estimators converge smoothly, but might find suboptimal solutions if the surrogate function is imprecise.
2267
Reposted by Marcel Hussing
Claas Voelcker @cvoelcker.bsky.social · 19/06/2025
Works that use #VAML/ #MuZero losses often use deterministic models. But if we want to use stochastic models to measure uncertainty or because we want to leverage current SOTA models such as #transformers and #diffusion, we need to take care! Naively translating the loss functions leads to mistakes!
174