Reposted by Yi-Hao Peng
STEPS unveils a selective on-policy self-distillation framework for math RL. By focusing on critical spans, it boosts average accuracy by 2.76% over existing methods, proving the value of targeted correction to enhance performance and cut training overhead. arxiv.org/abs/2605.10194
arxiv.org
STEPS: Selective On-Policy Self-Distillation for Reasoning
ArXiv link for STEPS: Selective On-Policy Self-Distillation for Reasoning