Or dig directly into our paper: openreview.net/forum?id=pPW...
openreview.net
Whatever Remains Must Be True: Filtering Drives Reasoning in LLMs...
Reinforcement Learning (RL) has become the _de facto_ standard for tuning LLMs to solve tasks involving reasoning.
However, growing evidence shows that models trained in such way often suffer from...