Can We Trust Item Response Theory for AI Evaluation?
arXiv paper · source date 2026-07-16 · added 2026-08-07 15:52:43 · updated 2026-08-07 15:52:43 · Open original blog
1
Problems / challenges / motivations
- IRT has quietly become the default aggregation method for serious independent evals (CAISI, ATLAS, GIM), imported from psychometrics.
- The AI regime violates the assumptions IRT was built for: few "test takers" (models), very many items, and non-normal ability distributions.
- Practitioners have no guidance on when the resulting ability estimates and item parameters can be trusted.
2
Key ideas
- Simulates 18,000 conditions across six LLM benchmarks to audit four estimator families: MML, MCMC, variational inference, and neural PSN.
- Classical estimators become computationally infeasible at benchmark scale.
- Scalable estimators produce unreliable item and ranking inferences when the model set is small or skewed — exactly the situation in frontier evaluation.
- Provides concrete sample-size and diagnostic guidance for when IRT is safe to use.
3
Why it matters for evals
- It is a timely, rigorous audit of a method the field adopted faster than it validated, and it arrives while government evaluators are actively building on IRT.
- The actionable output is the diagnostics: run them before publishing an IRT-aggregated capability claim.
- Two versions in under ten days signals active field interest.
- Caveat: a simulation-based preprint; conclusions are conditioned on the simulated benchmark structures and estimator implementations chosen.
Comments
No comments yet.