Anastasiia Pedan @pedanana.bsky.social · 25/08/2025my main takeaway from a talk on reward design in rl: ai only beat humans when they were asked not to collaborate 👀👀 120
Anastasiia Pedan @pedanana.bsky.social · 19/06/2025Would you be surprised to learn that many empirical implementations of value-aware model learning (VAML) algos, including MuZero, lead to incorrect model & value functions when training stochastic models 🤕? In our new @icmlconf.bsky.social 2025 paper, we show why this happens and how to fix it 🦾! 183