Excited to share our new paper on implicit safety alignment from crowd feedback that was accepted at ICML! We learn shared safety constraints from crowd feedback and show this can help alleviate reward misspecification. Kudos to my PhD student Qian Lin!
Paper: arxiv.org/abs/2605.21822
🧵1/7
arxiv.org
Implicit Safety Alignment from Crowd Preferences
Reinforcement Learning from Human Feedback (RLHF) can reveal implicit objectives such as safety considerations that go beyond task completion. In this work, we focus on the common safety criteria embe...