Reposted by Abigail Jacobs
Come find us at COLM, I’ll be presenting this paper next Thursday afternoon (10/8, oral session 6 and poster session 6). Full paper here: arxiv.org/abs/2609.08812
arxiv.org
What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks
Benchmarks play a central role in the development and governance of models, yet it is often unclear whether they actually measure the concepts they purport to measure (e.g., reasoning, refusal). We ad...