Reposted by André Cruz
LLM leaderboards aren't passive measurements — they're mechanisms that create incentives!
Why does the current evaluation protocol reward "benchmaxxing": post-training for the benchmark without improving model capability?
Our #ICML2026 paper gives a theoretical answer! 🧵