RL for real-world applications = offline learning + reward learning. How do we make this work?
Find out more at ICLR poster #377 at 10am today!
@gioramponi.bsky.social and I will be presenting our latest work on offline preference-based RL (joint w/ @gxxxr.bsky.social and Bernhard Schölkopf).