Across safe RL environments and a proof-of-concept LLM-style setting, our method reduces safety cost without explicit safety rewards, matches oracle safety baselines in task performance, and remains robust to preference imbalance.
Check out the paper here: arxiv.org/abs/2605.21822
🧵7/7
arxiv.org
Implicit Safety Alignment from Crowd Preferences
Reinforcement Learning from Human Feedback (RLHF) can reveal implicit objectives such as safety considerations that go beyond task completion. In this work, we focus on the common safety criteria embe...