Bo Liu (Benjamin Liu) @benjamin-eecs.bsky.social · 01/07/2025We're excited about self-play unlocking continuously improving agents. RL selects CoT patterns from LLMs. Games=perfect testing grounds. SPIRAL: models learn via self-competition. Kuhn Poker → +8.7% math, +18.1 Minerva Math! 🃏 Paper: huggingface.co/papers/2506.... Code: github.com/spiral-rl/spiral 2175
Reposted by Bo Liu (Benjamin Liu)Ksenia Se / Turing Post @turingpost.bsky.social · 29/11/2024Natural Language Reinforcement Learning (NLRL) redefines Reinforcement Learning (RL). NLRL's main idea: The core parts of RL like goals, strategies, and evaluation methods are reimagined using natural language instead of rigid math. Let's explore this approach more precisely🧵 121