A bit outside my usual territory, but the problem is very cool and with lot of interesting open questions. Worth a look if you care about the theoretical underpinnings of RL for pass@k objectives.
arxiv.org/abs/2607.17823
arxiv.org
Theoretical Foundations of $\max$@$k$ Reinforcement Learning
Reinforcement Learning is a cornerstone technique for modern large reasoning models. Usually, for difficult tasks such as code generation and theorem proving, the agent is evaluated by generating $K$ ...