🔥Excited to share our new work: "A Practitioner's Guide to Multi-turn Agentic Reinforcement Learning"!
We study what actually works for agentic multi-turn RL with varying 🌎Environment, 🤖Policy, and ⭐Reward.
We conduct various ablations and empirical analysis on 🧩TextWorld, 🧙ALFWorld, and 🧑💻SWE-Gym.