Sign in

Daniel Brown

@daniel-brown.bsky.social
586 followers 13 following 51 posts

CS assistant prof @Utah. Researches human-robot interaction, human-in-the-loop ML, AI safety and alignment. users.cs.utah.edu/~dsbrown

PostsRepliesMedia
Daniel Brown @daniel-brown.bsky.social · 03/06/2026
Across safe RL environments and a proof-of-concept LLM-style setting, our method reduces safety cost without explicit safety rewards, matches oracle safety baselines in task performance, and remains robust to preference imbalance. Check out the paper here: arxiv.org/abs/2605.21822 🧵7/7
arxiv.org
Implicit Safety Alignment from Crowd Preferences
Reinforcement Learning from Human Feedback (RLHF) can reveal implicit objectives such as safety considerations that go beyond task completion. In this work, we focus on the common safety criteria embe...
000
Daniel Brown @daniel-brown.bsky.social · 03/06/2026
Idea: compose safe skills instead of combining safety rewards. First learn preference-aligned skills that capture different user preferences from crowd data. Then compose these skills to solve new downstream tasks. The high-level policy solves the task. The low-level skills provide safety. 🧵6/7
100
Daniel Brown @daniel-brown.bsky.social · 03/06/2026
Why? The reward learned from crowd preferences inevitably entangles shared safety criteria with user-specific preferences. When preference data are imbalanced, it can overfit to the majority user’s objective, hurting downstream task performance and making the reward trade-off hard to tune. 🧵5/7
100
Daniel Brown @daniel-brown.bsky.social · 03/06/2026
A natural solution is reward combination: learn a reward from crowd preferences, then combine it with the downstream task reward. But we show this approach can be fragile. 🧵4/7
100
Daniel Brown @daniel-brown.bsky.social · 03/06/2026
Our key intuition: crowd preferences may contain diverse user goals, but shared safety criteria. Different users may prefer different trajectories for their own purposes, yet consistently avoid unsafe regions or actions. 🧵3/7
100
Daniel Brown @daniel-brown.bsky.social · 03/06/2026
Motivation: task rewards are often incomplete. A reward may tell an agent to reach a goal, but not tell it which states or actions are unsafe. Optimizing such a safety-agnostic reward can lead to high task reward but unsafe behavior. 🧵2/7
100
Daniel Brown @daniel-brown.bsky.social · 03/06/2026
Excited to share our new paper on implicit safety alignment from crowd feedback that was accepted at ICML! We learn shared safety constraints from crowd feedback and show this can help alleviate reward misspecification. Kudos to my PhD student Qian Lin! Paper: arxiv.org/abs/2605.21822 🧵1/7
arxiv.org
Implicit Safety Alignment from Crowd Preferences
Reinforcement Learning from Human Feedback (RLHF) can reveal implicit objectives such as safety considerations that go beyond task completion. In this work, we focus on the common safety criteria embe...
130
Daniel Brown @daniel-brown.bsky.social · 03/06/2026
This motivates per-user adaptation and personalization, an exciting area of future work! This has been a really fun collaboration within the Utah Robotics Center. Kudos to co-authors: Colin Rubow, Eric Brewer, Ian Bales, and Haohan Zhang! arxiv.org/abs/2603.06779 🧵8/8
arxiv.org
A Multi-Layer Sim-to-Real Framework for Gaze-Driven Assistive Neck Exoskeletons
Dropped head syndrome, caused by neck muscle weakness from neurological diseases, severely impairs an individual's ability to support and move their head, causing pain and making everyday tasks challe...
000
Daniel Brown @daniel-brown.bsky.social · 03/06/2026
We find that our two of our three data-driven gaze controllers are competitive or superior to prior state-of-the-art. Interestingly, there is no single overall "best" controller. We found three non-dominated controllers and which one is best depends on the individual user. 🧵7/8
100
Daniel Brown @daniel-brown.bsky.social · 03/06/2026
We first evaluate the controllers in VR by having the environment respond to eye positions to rotate around the users head to imitate head motions. This lets us quickly reject bad controllers before physical deployment. 🧵6/8
100
Daniel Brown @daniel-brown.bsky.social · 03/06/2026
We developed data driven controllers by collecting training data (coupled eye and head movements from healthy individuals) using virtual reality (VR). VR also gives us a way to develop a digital twin of the physical exoskeleton. 🧵5/8
100
Daniel Brown @daniel-brown.bsky.social · 03/06/2026
We want to develop controllers that are natural and intuitive and also be able to safely test them before use with a physical robot attached to the head. In our recent ICRA paper, we present the utilization of VR to develop and test new data driven controllers. 🧵4/8
100
Daniel Brown @daniel-brown.bsky.social · 03/06/2026
But what should the software be to help users regain natural neck motion? Haohan and I have an NIH Trailblazer dedicated to studying this question. To control the braces we proposed to develop novel head-neck controllers that track a user's eye gaze. 🧵3/8
100
Daniel Brown @daniel-brown.bsky.social · 03/06/2026
Our ultimate goal is to develop robotic neck exoskeletons to restore head and neck motion to those with diseases that weaken the neck (e.g, ALS). Prof. Haohan Zhang and his lab at Utah have developed some of the world's first powered neck exoskeletons. 🧵2/8
100
Daniel Brown @daniel-brown.bsky.social · 03/06/2026
Our paper "A Multi-Layer Sim-to-Real Framework for Gaze-Driven Assistive Neck Exoskeletons" is being presented this week at ICRA 2026! arxiv.org/abs/2603.06779 🧵1/8
100
Daniel Brown @daniel-brown.bsky.social · 19/05/2026
This work offers a new mechanistic lens for understanding why normalization, resets, and classification-based value learning work in deep RL — and provides practical guidance for choosing among them depending on reward sparsity. Full paper: openreview.net/forum?id=VNV... 7/N
openreview.net
Understanding the Effects of Neuron Dominance in Deep Reinforcement...
Recent studies in deep reinforcement learning have revealed that neural networks tend to lose their capacity to adapt to new targets over the course of training. The proliferation of inactive...
010
Daniel Brown @daniel-brown.bsky.social · 19/05/2026
4. We provide a novel theoretical explanation for why classification losses like HL-Gauss resist representation collapse: unlike MSE regression, the softmax nonlinearity prevents gradients from vanishing even when the representation collapses to zero. 6/N
openreview.net
Understanding the Effects of Neuron Dominance in Deep Reinforcement...
Recent studies in deep reinforcement learning have revealed that neural networks tend to lose their capacity to adapt to new targets over the course of training. The proliferation of inactive...
110
Daniel Brown @daniel-brown.bsky.social · 19/05/2026
3. We evaluate five mitigation strategies (Reset, ReDo, Pruning, HL-Gauss, and LayerNorm) and found that approaches which keep peak activations low tend to have higher representational capacity, fewer dormant neurons, and better performance. 5/N
openreview.net
Understanding the Effects of Neuron Dominance in Deep Reinforcement...
Recent studies in deep reinforcement learning have revealed that neural networks tend to lose their capacity to adapt to new targets over the course of training. The proliferation of inactive...
100
Daniel Brown @daniel-brown.bsky.social · 19/05/2026
2. We connect neuron dominance to representation collapse — showing that higher peak activations correlate with lower effective rank in the feature matrix, meaning the network loses the ability to distinguish between inputs. 4/N
openreview.net
Understanding the Effects of Neuron Dominance in Deep Reinforcement...
Recent studies in deep reinforcement learning have revealed that neural networks tend to lose their capacity to adapt to new targets over the course of training. The proliferation of inactive...
100
Daniel Brown @daniel-brown.bsky.social · 19/05/2026
Key findings: 1. Dominant neurons reliably co-occur with complete saturation of the subsequent layer across multiple sparse-reward visual control tasks, effectively killing the network's ability to learn anything new. 3/N
openreview.net
Understanding the Effects of Neuron Dominance in Deep Reinforcement...
Recent studies in deep reinforcement learning have revealed that neural networks tend to lose their capacity to adapt to new targets over the course of training. The proliferation of inactive...
100
Daniel Brown @daniel-brown.bsky.social · 19/05/2026
Deep RL networks lose plasticity — the ability to update their predictions based on new experience. We identify dominant neurons (neurons with disproportionately large activations) as a key culprit. 2/N
100
Daniel Brown @daniel-brown.bsky.social · 19/05/2026
Super excited about our lab's new paper led by my PhD student Zifan Wu: "Understanding the Effects of Neuron Dominance in Deep Reinforcement Learning", that has been published in Transactions on Machine Learning Research (TMLR)! openreview.net/forum?id=VNV... 1/N
openreview.net
Understanding the Effects of Neuron Dominance in Deep Reinforcement...
Recent studies in deep reinforcement learning have revealed that neural networks tend to lose their capacity to adapt to new targets over the course of training. The proliferation of inactive...
111
Daniel Brown @daniel-brown.bsky.social · 23/03/2026
This work was led by my PhD student Connor Mattson with Varun Raveendra, Ellen Novoseller, Nicholas Waytowich, and Vernon J. Lawhern. 10/N
001
Daniel Brown @daniel-brown.bsky.social · 23/03/2026
Our paper highlights a new method for teaching robots that is designed around what humans can feasibly demonstrate. Lots of exciting follow up work to come in this area! 9/N
101
Daniel Brown @daniel-brown.bsky.social · 23/03/2026
- On physical robots, R2BC outperforms centralized BC by 3.25x and 5.9x on two cooperative tasks. - The variability introduced by autonomous teammates during training acts as a natural form of data augmentation, making R2BC policies more robust. 8/N
101
Daniel Brown @daniel-brown.bsky.social · 23/03/2026
Some of our exciting findings indicate that: - R2BC matches or outperforms a privileged joint-action behavior cloning baseline across four simulated tasks, despite never seeing a single centralized demonstration. 7/N
101
Daniel Brown @daniel-brown.bsky.social · 23/03/2026
Not only does this work, our method is simpler and better than giving an "oracle" demonstrator control of the whole team simultaneously in an offline setting. 6/N
101
Daniel Brown @daniel-brown.bsky.social · 23/03/2026
In our method, Round-Robin Behavior Cloning (R2BC), the human controls one agent while the others run their current learned policies. Then, the demonstrator rotates and teaches the next robot. This process repeats while agents periodically update their policies with their demos. 5/N
101
Daniel Brown @daniel-brown.bsky.social · 23/03/2026
We tackled the following question: How can we teach robot teams if a human can only provide demonstrations to one robot at a time? 4/N
101
Daniel Brown @daniel-brown.bsky.social · 23/03/2026
Problem: Teaching a team of robots by demonstration is hard. A single human cannot teleoperate many robots at once and expect any of them to learn something useful. However, prior work has unrealistically assumed that a single human can provide these joint demos to agents! 3/N
101
Daniel Brown @daniel-brown.bsky.social · 23/03/2026
R2BC: Multi-Agent Imitation Learning from Single-Agent Demonstrations Full Paper: arxiv.org/abs/2510.18085 Code, and Videos: sites.google.com/view/r2bc/home 2/N
arxiv.org
R2BC: Multi-Agent Imitation Learning from Single-Agent Demonstrations
Imitation Learning (IL) is a natural way for humans to teach robots, particularly when high-quality demonstrations are easy to obtain. While IL has been widely applied to single-robot settings, relati...
101
Daniel Brown @daniel-brown.bsky.social · 23/03/2026
How can you do imitation learning for multi-agent systems if the demonstrator is just a single human? Unless you're a doctor octopus genius, this is really hard! We study this problem in a new paper that will be at ICRA'26! 1/N
130
Daniel Brown @daniel-brown.bsky.social · 10/11/2025
The Kahlert School of Computing at the University of Utah is hiring for multiple faculty positions! We're especially interested in growing in the areas of human-centered AI and robotics! utah.peopleadmin.com/postings/190...
utah.peopleadmin.com
Assistant Professor for The Kahlert School of Computing
000
Daniel Brown @daniel-brown.bsky.social · 10/10/2025
We hope this work can help inspire the development of better AI alignment tests and evaluations for LLM reward models. Check out the workshop paper here: anamarasovic.com/publications... 8/8
anamarasovic.com
030
Daniel Brown @daniel-brown.bsky.social · 10/10/2025
We applied this approach to RewardBench and found evidence that much of the data in safety and reasoning datasets may be redundant (44% for safety and 24% for reasoning) and that this can lead to inflated alignment scores. 7/8
100
Daniel Brown @daniel-brown.bsky.social · 10/10/2025
By scaling up these ideas to LLMs, we can now estimate the set of reward model weights (weights that map the last decoder hidden state to a scalar output) that are consistent with a preference alignment dataset and also identify redundant and non-redundant examples in the preference dataset. 6/8
100
Daniel Brown @daniel-brown.bsky.social · 10/10/2025
Once you find these core demonstrations or comparisons you can use them to craft efficient alignment tests. But until recently, we were only able to empirically test these ideas on simple toy domains. 5/8
100
Daniel Brown @daniel-brown.bsky.social · 10/10/2025
The main idea was that for linear rewards, we can determine, via an intersection of half-spaces, the set of reward functions that make a policy optimal and that this set of rewards is defined by a small number of "non-redundant" demonstrations or comparisons. 4/8
100
Daniel Brown @daniel-brown.bsky.social · 10/10/2025
It was a fun paper and has some interesting nuggets, like the fact that there exist sufficient conditions under which we can verify exact and approximate AI alignment across an infinite set of deployment environments via a constant-query-complexity test. 3/8
120
Daniel Brown @daniel-brown.bsky.social · 10/10/2025
As some background, a couple of years ago I worked with Jordan Schneider, @scottniekum.bsky.social, and Anca Dragan on what we called "Value Alignment Verification" with the goal of efficiently testing whether an AI system is aligned with human values. arxiv.org/abs/2012.01557 2/8
arxiv.org
Value Alignment Verification
As humans interact with autonomous agents to perform increasingly complicated, potentially risky tasks, it is important to be able to efficiently evaluate an agent's performance and correctness. In th...
120
Daniel Brown @daniel-brown.bsky.social · 10/10/2025
Can you trust your reward model alignment scores? New work presented today at the COLM Workshop on Socially Responsible Language Modelling Research led by Purbid Bambroo and in collaboration with @anamarasovic.bsky.social that probes LLM preference test sets for redundancy and inflated scores. 1/8
121
Daniel Brown @daniel-brown.bsky.social · 29/09/2025
This was a really fun collaboration with Jordan Thompson, Britton Jordan, and Alan Kuntz. Check out our paper here: openreview.net/forum?id=K7K... 5/5
openreview.net
Agreement Volatility: A Second-Order Metric for Uncertainty...
Autonomous surgical robots are a promising solution to the increasing demand for surgery amid a shortage of surgeons. Recent work has proposed learning-based approaches for the autonomous...
000
Daniel Brown @daniel-brown.bsky.social · 29/09/2025
Our approach also enables uncertainty attribution! We can backpropagate uncertainty estimates into an input point cloud to visualize and interpret the robot's uncertainty. If you're at #CoRL25, check out Jordan Thompson's talk and poster (Spotlight 6 & Poster 3). 4/5
100
Daniel Brown @daniel-brown.bsky.social · 29/09/2025
We apply our approach to surgically-inspired deformable tissue manipulation and find it achieves a 10% lower reliance on human interventions compared to prior work that leverages variance-based uncertainty estimates. 3/5
100
Daniel Brown @daniel-brown.bsky.social · 29/09/2025
Inspired by prior work on active, uncertainty-aware human-robot hand-offs like Ryan Hoque and @ken-goldberg.bsky.social's ThriftyDAgger (arxiv.org/abs/2109.08273), we show that agreement volatility enables robots to know when they need help so they can request appropriate human interventions. 2/5
100
Daniel Brown @daniel-brown.bsky.social · 29/09/2025
Check out our new paper being presented today at #CoRL2025 on uncertainty quantification: openreview.net/forum?id=K7K.... We propose a new second-order metric for uncertainty quantification in robot learning that we call "Agreement Volatility." 1/5
100
Daniel Brown @daniel-brown.bsky.social · 25/09/2025
Excited to announce that my lab's research was recently highlighted in an AI Magazine article: onlinelibrary.wiley.com/doi/pdf/10.1...
onlinelibrary.wiley.com
Toward robust, interactive, and human‐aligned AI systems
Ensuring that AI systems do what we, as humans, actually want them to do is one of the biggest open research challenges in AI alignment and safety. My research seeks to directly address this challeng...
000
Daniel Brown @daniel-brown.bsky.social · 03/03/2025
If you're in Melbourne, come check out Connor's talk in the Teleoperation and Shared Control session today! Paper: arxiv.org/abs/2501.08389 Website: sites.google.com/view/zerosho... This is joint work with two of my other amazing PhD students Zohre Karimi and Atharv Belsare! 3/3
000
Daniel Brown @daniel-brown.bsky.social · 03/03/2025
We study how to enable robots to use end-effector vision to estimate zero-shot human intents in conjunction with blended control to help humans accomplish manipulation tasks like grocery shelving with unknown and dynamically changing object locations. 2/3
100
Daniel Brown @daniel-brown.bsky.social · 03/03/2025
Shared autonomy systems have been around for a long time but most approaches require a learned or specified set of possible human goals or intents. I'm excited for my student Connor Mattson to present our work at #HRI2025 on a zero-shot, vision-only shared autonomy (VOSA) framework. 1/3
110