Sign in

tylerkastner.bsky.social

@tylerkastner.bsky.social
5 followers 21 following 0 posts
PostsRepliesMedia
Reposted by @tylerkastner.bsky.social
Claas Voelcker @cvoelcker.bsky.social · 11/02/2025
Do you want to get the most out of your samples, but increasing the update steps just destabilizes RL training? Our #ICLR2025 spotlight 🎉 paper shows that using the values of unseen actions causes instability in continuous state-action domains and how to combat this problem with learned models!
3207