Sign in

Marcel Hussing

@marcelhussing.bsky.social
3K followers 336 following 215 posts

PhD student at the University of Pennsylvania. Prev, intern at MSR, and Meta FAIR. Interested in reliable and replicable reinforcement learning, robotics and knowledge discovery: marcelhussing.github.io All posts are my own.

PostsRepliesMedia
Marcel Hussing @marcelhussing.bsky.social · 07/10/2026
Like many, I've thought a lot about the developments in theoretical research. And while I'm not as confident as others here on what to make of them, I think one thing is clear. Recent results are yet another historical marker likely left by RL and PG-style methods. #math #openai
010
Reposted by Marcel Hussing
Claas Voelcker @cvoelcker.bsky.social · 01/10/2026
⏰ More time to submit! ⏰ We hope your drafts are getting into shape and we got good news! The #oopsiedata workshop @ #CoRL2026 🤖 deadline is now October 5th. Polish up your work and send it our way! We are looking forward to all the excellent submissions!
122
Marcel Hussing @marcelhussing.bsky.social · 03/09/2026
To find out more about our other efforts surrounding failure and suboptimal datasets, check out the Oopsie-Data project at oopsie-data.com
oopsie-data.com
Home
Tools for collecting, annotating, and managing robotic manipulation failure data — part of the Oopsie Data project.
000
Marcel Hussing @marcelhussing.bsky.social · 03/09/2026
This workshop is co-organized with the Oopsie-Data coordination team: @cvoelcker.bsky.social , Siddhant Agarwal, Arpit Bahety, Zhiyuan (Paul) Zhou, @renwang435.bsky.social, Jiahui Chen, Maria Attarian, Carl Qi, @maxbrudolph.bsky.social, Roberto Martín-Martín, Sergey Levine, Peter Stone, Amy Zhang
100
Marcel Hussing @marcelhussing.bsky.social · 03/09/2026
🤖 Coming to CoRL 2026: The first workshop on imperfect robotics data! We’ll explore how to collect, curate & use failures 💢, OOD events❓, and unexpected human behavior 🤔 - data every lab produces but rarely documents, shares, or reuses. Check out workshop.oopsie-data.com
142
Marcel Hussing @marcelhussing.bsky.social · 01/08/2026
I've tried to be a good reviewer, read all the papers carefully, and formed my own view. I'm getting page long, clearly LLM generated rebuttals. It is not my job to prompt your LLM into convincing me. What are people's solutions to this? Should I message the AC? #NeurIPS2026
010
Reposted by Marcel Hussing
Claas Voelcker @cvoelcker.bsky.social · 30/07/2026
This is a robot failing to grasp a ball. Almost every robot lab produces clips like this daily… and almost all of them get thrown away. This is the most abundant but underused resource in robot learning. We’re collecting all of it now as “OopsieData”, please join us at oopsie-data.com! (1/13)
17413
Marcel Hussing @marcelhussing.bsky.social · 23/07/2026
At this point I'm not even sure... maybe I prefer LLM reviews.
040
Marcel Hussing @marcelhussing.bsky.social · 22/07/2026
It's this time of year
static.klipy.com
Mr Bean: Rowan Atkinson Waiting
ALT: Mr Bean: Rowan Atkinson Waiting
010
Marcel Hussing @marcelhussing.bsky.social · 29/06/2026
Post is a bit delayed but last week I passed my dissertation defense on algorithm output stability in RL. 🎉
140
Marcel Hussing @marcelhussing.bsky.social · 26/05/2026
This was a very fun project. In Behavior-Consistent Deep RL, we provide a method that aligns the behavior of independently trained policies. It turns out, this works even in high dimensional spaces. Here are 6 seeds of Humanoids (all ca same return). (left) Baseline (right) Ours.
092
Marcel Hussing @marcelhussing.bsky.social · 22/05/2026
This was a fun collaboration between Penn, Princeton, and UT Austin. 🔗 arxiv.org/abs/2605.21214 🧑‍🎓 @liv-daliberti.bsky.social @cvoelcker.bsky.social @ben-eysenbach.bsky.social @ericeaton.bsky.social
arxiv.org
Behavior-Consistent Deep Reinforcement Learning
Reinforcement learning (RL) often exhibits high variance across training runs, leading to unreliable performance and posing a major challenge to deployment in real-world domains. In this work, we addr...
020
Marcel Hussing @marcelhussing.bsky.social · 22/05/2026
On standard benchmarks, QED reduces policy divergence across independent runs by up to two orders of magnitude while maintaining competitive performance.
110
Marcel Hussing @marcelhussing.bsky.social · 22/05/2026
This motivates Q-value Expectile Disagreement, or QED. QED uses critic disagreement as a proxy for likely behavioral divergence. When disagreement is high, QED increases temperature and keeps the policy exploratory and consistent. When disagreement is low, the policy can specialize.
120
Marcel Hussing @marcelhussing.bsky.social · 22/05/2026
But simply increasing entropy is not enough. Too much entropy can hurt optimization and amplify off-policy error, especially when the critic is queried on actions that are poorly supported by the data.
120
Marcel Hussing @marcelhussing.bsky.social · 22/05/2026
Our main observation is that maximum-entropy RL already gives us a mechanism for controlling this divergence. Entropy regularization anchors policies toward a shared prior, which can reduce the space of behaviors discovered across independent runs.
120
Marcel Hussing @marcelhussing.bsky.social · 22/05/2026
The core challenge is that, during training, two independent runs cannot directly compare their learned policies. Each run only has access to its own data, critics, and policy. So we need a signal within a single run that predicts when executions may diverge.
110
Marcel Hussing @marcelhussing.bsky.social · 22/05/2026
In Behavior-Consistent Deep RL, we study consistency in continuous state-action spaces, such as those that arise in robotics. The goal is to make independent training runs produce behaviorally similar policies, not just high-performing ones.
110
Marcel Hussing @marcelhussing.bsky.social · 22/05/2026
RL algorithms are notoriously brittle. Even when we use the same algorithm in the same environment, independent runs can produce drastically different behavior. This makes it hard to audit, compare, and reliably update learned policies.
110
Marcel Hussing @marcelhussing.bsky.social · 22/05/2026
In a standard ML deployment pipeline, new data is regularly added, a model is retrained, and the updated model is deployed. For this to be useful, updates need to be auditable, reliable, and consistent: we want improvements from new data without unexpected new failure modes.
120
Marcel Hussing @marcelhussing.bsky.social · 22/05/2026
🚨 New Preprint Alert: Behavior-Consistent Deep Reinforcement Learning 🚨 TLDR: We introduce an approach that achieves behavioral similarity across independent algorithm executions in continuous state-action space deep RL.
1294
Marcel Hussing @marcelhussing.bsky.social · 27/04/2026
I guess I phrased it a bit awkwardly in the hectic of presenting. It's not about the number 27 specifically but rather about the fact that even though these models ostensibly have different architectures and they are highly non convex, they still predict the **same number**.
010
Marcel Hussing @marcelhussing.bsky.social · 27/04/2026
bsky.app/profile/marc...
100
Marcel Hussing @marcelhussing.bsky.social · 27/04/2026
In our work we show that for certain model classes (including relu neural networks) loss minimization (not global optimality) implies agreement in prediction space. If you are interested in how we get to that result @aaroth.bsky.social wrote a summary thread here bsky.app/profile/aaro...
140
Marcel Hussing @marcelhussing.bsky.social · 27/04/2026
Unclear how this is affected by reasoning and post training but this was a kind of a surprising thing that came up a few years ago. Here is a reddit thread www.reddit.com/r/dataisbeau...
reddit.com
From the dataisbeautiful community on Reddit: LLMs and the number 27: Myth tested with 800 prompts [OC]
Explore this post and more from the dataisbeautiful community
120
Marcel Hussing @marcelhussing.bsky.social · 26/04/2026
Why do all LLMs predict 27 as their favorite number? There may be a principled explanation. Learn more at Agents in the Wild at #ICLR2026. @ericeaton.bsky.social, me, @surbhigoel.bsky.social, @mkearnsphilly.bsky.social, @aaroth.bsky.social, @sikatasengupta.bsky.social, @optimistsinc.bsky.social
2122
Marcel Hussing @marcelhussing.bsky.social · 25/04/2026
Time for round number two. Stop by #4406 to learn about stability guarantees in RL. #ICLR2026 @ericeaton.bsky.social @mkearnsphilly.bsky.social @aaroth.bsky.social @sikatasengupta.bsky.social @optimistsinc.bsky.social
092
Marcel Hussing @marcelhussing.bsky.social · 25/04/2026
At #ICLR2026 presenting our first poster in the morning on Relative Entropy Pathwise Policy Optimization. Stop by at #4613. 🧑‍🎓 @cvoelcker.bsky.social, @axelbrunnbauer.bsky.social, Michal Naumann, Pieter Abbeel, @ericeaton.bsky.social, Radu Grosu, @sologen.bsky.social @igilitschenski.bsky.social
0164
Reposted by Marcel Hussing
JHU Computer Science @jhucompsci.bsky.social · 21/04/2026
In “Replicable Reinforcement Learning with Linear Function Approximation,” @optimistsinc.bsky.social, @marcelhussing.bsky.social, @mkearnsphilly.bsky.social, @aaroth.bsky.social, @sikatasengupta.bsky.social, & more develop replicable methods for linear function approximation in RL: (5/12)
arxiv.org
Replicable Reinforcement Learning with Linear Function Approximation
Replication of experimental results has been a challenge faced by many scientific disciplines, including the field of machine learning. Recent work on the theory of machine learning has formalized rep...
173
Marcel Hussing @marcelhussing.bsky.social · 20/04/2026
Greetings from Rio, killing some time before #ICLR2026 @iclr-conf.bsky.social
040
Marcel Hussing @marcelhussing.bsky.social · 17/04/2026
Getting ready to fly to Rio. 😎🇧🇷
000
Marcel Hussing @marcelhussing.bsky.social · 03/04/2026
Why is the default option in the rebuttal acknowledgements at ICML to accept the paper? I'm very confused on how to use these buttons. Can someone explain them to me?
020
Marcel Hussing @marcelhussing.bsky.social · 25/03/2026
That is even if the baseline reported 80% on a similar task
030
Reposted by Marcel Hussing
Aaron Roth @aaroth.bsky.social · 12/03/2026
Neural networks are highly non-convex, so approximate error minimizers need not look anything like each other in parameter space. But we show that nevertheless (for many model sizes) approximate error minimizers must closely agree in function/prediction space despite this!
1107
Marcel Hussing @marcelhussing.bsky.social · 10/03/2026
It's this time of the year again: your baselines cannot be PPO and SAC.
121
Reposted by Marcel Hussing
Aaron Roth @aaroth.bsky.social · 09/03/2026
Michael @mkearnsphilly.bsky.social ) and I wrote a blog post about our experiences using AI for research, and our thoughts on what these developments will mean for research, publication, and education: www.amazon.science/blog/how-ai-...
amazon.science
How AI is changing the nature of mathematical research
What machine learning theorists learned using AI agents to generate proofs — and what comes next.
13013
Marcel Hussing @marcelhussing.bsky.social · 08/03/2026
I'm so glad that so many research problems are finally being treated as first class citizens rather than afterthoughts. 🤔
010
Marcel Hussing @marcelhussing.bsky.social · 06/03/2026
Too many papers sound like this Hierarchical Context-Aware Diffusion-Transformer Meta-World-Model Reinforcement Learning with Causally Disentangled Preference-Aligned Self-Supervised Compositional Multi-Scale Latent Skill Priors for Long-Horizon Generalist Decision Making
0110
Marcel Hussing @marcelhussing.bsky.social · 05/03/2026
One reason I work on replicable and consistent RL is because it is has always been at the top of the list of criteria for reliability.
020
Marcel Hussing @marcelhussing.bsky.social · 03/03/2026
Excuse me? Surely telling me that didn't require much thinking.
210
Marcel Hussing @marcelhussing.bsky.social · 03/03/2026
I have seen multiple times now that a reviewer said sth like: the proofs are simple -> reject the paper. That is completely counter-productive. A theorem needs to generate new insights. If we learn something new from something simple that should be preferred. Don't believe me? Ask someone famous:
010
Marcel Hussing @marcelhussing.bsky.social · 26/02/2026
Why are you doing this to me
010
Marcel Hussing @marcelhussing.bsky.social · 16/02/2026
I think it's relatively simple. One side already has the job they want and the other side needs citations to get that job. And everyone tells me advertising work is how you get citations. I have been tempted to go back because on bsky, interactions have become fewer and fewer. Not going to though...
060
Marcel Hussing @marcelhussing.bsky.social · 15/02/2026
Yet somehow every now and then a paper becomes very popular even though its findings are similar to those of many others. This paper gets cited while the others don't. Was it just luck?
010
Marcel Hussing @marcelhussing.bsky.social · 15/02/2026
While I agree with the sentiment let me play devil's advocate. You only get invited to give talks if your work is already well known. Conferences have become too large to even find relevant people. And social media posts are only marginally relevant if you don't already have a large following.
220
Marcel Hussing @marcelhussing.bsky.social · 14/02/2026
I understand that that is okay but at some point it honestly becomes disheartening if you constantly have to reach out to people. There are others who don't seem to have to; what are they doing differently?
100
Marcel Hussing @marcelhussing.bsky.social · 14/02/2026
I’m trying to understand whether this is mostly about keyword mismatch, venue visibility, social media, etc. For example, when I search terms like “high update ratio RL” on Scholar, our papers show up near the top. scholar.google.com/scholar?hl=e... Where are things going wrong?
scholar.google.com
110
Marcel Hussing @marcelhussing.bsky.social · 14/02/2026
I’ve been thinking about a practical question and would love some opinions: How do your papers actually get discovered/cited? I was searching for recent work on high update ratio RL and found several very closely related papers tackling the same failure modes we study. None cited our earlier work.
592
Reposted by Marcel Hussing
Igor Gilitschenski @igilitschenski.bsky.social · 13/02/2026
🚀 Excited to share REPPO, a new on-policy RL agent! TL;DR: Replace PPO with REPPO for fewer hyperparameter headaches and more robust training. REPPO, led by @cvoelcker.bsky.social, will be presented at ICLR 2026. How does it work? 🧵👇
12510
Reposted by Marcel Hussing
Lukas Heinrich @lukasheinrich.com · 01/02/2026
Scaling Laws in Particle Physics Data! This is a result I've been itching to share and it's finally out. One of the big open questions is how much better AI-based methods at particle colliders can still become. 1/4
3167