Reposted by Shyamgopal Karthik
New Paper: arxiv.org/abs/2606.19370
Self-play yields capabilities but requires frustrating cost-function tuning. Surprisingly, just 30 minutes of demonstration data produces much more human-like driving policies!
Led by @daphne-cornelisse.bsky.social
Website: spiced-self-play.com