Coding agents are evaluated with unit tests:
more tests passed = better model.
But if tests or feedback are accessible, models may learn to game them.
We introduce CapCode to detect suspiciously high scores, and CapReward to discourage them during RL.
🧵1/10