Sign in

Ryan Steed

@rbsteed.com
571 followers 167 following 24 posts

AI Policy Fellow @ Princeton | PhD Carnegie Mellon | privacy, accountability, & algorithmic systems

PostsRepliesMedia
Ryan Steed @rbsteed.com · 19/02/2026
GLMMs have other benefits, too: - We can estimate question difficulties to identify problematic questions and other patterns in benchmarks. - Variance decomposition (between- and within-questions) can highlight nuances in performance between tasks, languages, and other subsets of a benchmark.
Fig. 6a from the paper: Distribution of estimated question difficulties by domain and labeled difficulty (GPQA-Diamond). Distribution of random effects in each domain. Chemistry questions were particularly difficult
for the LLMs we tested. Each dot indicates a GPQA-Diamond question’s GLMM-estimated difficulty (i.e., random effect value). Box plots display quartiles and violin plots display estimated density. These estimates show that GPQA-Diamond’s chemistry questions were particularly difficult for the 22 tested LLMs. On the other hand, question difficulty for LLMs has a weak relationship with question-writer-labeled difficulty. This may suggest that humans and the tested LLMs find different questions difficult, and/or could call into question whether writer annotations are accurate even for human difficulty.Fig. 6b from the paper: Distribution of estimated question difficulties by domain and labeled difficulty (GPQA-Diamond). Distribution of random effects for the 191 questions at the three most common writer-annotated difficulty levels. Question difficulty for LLMs has a weak relationship with
human-labeled difficulty. Each dot indicates a GPQA-Diamond question’s GLMM-estimated difficulty (i.e., random effect value). Box plots display quartiles and violin plots display estimated density. These estimates show that GPQA-Diamond’s chemistry questions were particularly difficult for the 22 tested LLMs. On the other hand, question difficulty for LLMs has a weak relationship with question-writer-labeled difficulty. This may suggest that humans and the tested LLMs find different questions difficult, and/or could call into question whether writer annotations are accurate even for human difficulty.
100
Ryan Steed @rbsteed.com · 19/02/2026
AI evals rarely specify which question is being answered — but the choice matters, especially when it comes to computing error bars. (Assuming error bars are included at all…) In particular, error bars for generalized accuracy tend to be larger and may yield different rankings.
Fig. 1 from the paper: Comparing accuracy estimates (GPQA-Diamond). Lower plots show the estimated accuracy of a selection of tested LLMs* with 95% confidence intervals. Upper plots show corresponding confidence interval (CI) widths. Generalized accuracy CIs are larger than benchmark accuracy CIs because they account for the selection of benchmark items from a superpopulation. Notably, some pairs of LLMs may have significantly different benchmark accuracy but not generalized accuracy. The simple average (pink) estimates reflect the average across all n benchmark questions with standard error calculated as standard deviation of results divided by √n. For estimates of benchmark accuracy, the simple average method results in under-confident CIs compared to a valid regression-free method (blue). For estimates of generalized accuracy, the simple average method provides valid CIs, but precision can be increased by running more trials per item (as in the regression-free method). Generalized linear mixed model (GLMM, orange) estimates require additional assumptions but further increase precision.
100
Ryan Steed @rbsteed.com · 11/02/2026
Automated benchmarks are not all you need, but they are popular tools in AI development. Hoping this doc is a foundation for future guidelines on field testing and other kinds of evals.
Table I.1 from https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.800-2.ipd.pdf
100
Ryan Steed @rbsteed.com · 11/02/2026
Section 3 covers critical practices related to responsible and transparent reporting — including uncertainty quantification, reproducibility, and properly qualified claims.
Table 3.1 from https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.800-2.ipd.pdf
100
Ryan Steed @rbsteed.com · 11/02/2026
I’m especially excited about the focus on practical measurement validity. Sections 1 describes ways to assess the relationship between the contents of a benchmark and what evaluators really want to measure.
Table 1.1 from https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.800-2.ipd.pdf
100