GLMMs have other benefits, too:
- We can estimate question difficulties to identify problematic questions and other patterns in benchmarks.
- Variance decomposition (between- and within-questions) can highlight nuances in performance between tasks, languages, and other subsets of a benchmark.