AI & Agent Evaluation
2,151total visitsadmin

Can We Trust Item Response Theory for AI Evaluation?

arXiv paper · source date 2026-07-16 · added 2026-08-07 15:52:43 · updated 2026-08-07 15:52:43 · Open original blog

Problems / challenges / motivations

  • IRT has quietly become the default aggregation method for serious independent evals (CAISI, ATLAS, GIM), imported from psychometrics.
  • The AI regime violates the assumptions IRT was built for: few "test takers" (models), very many items, and non-normal ability distributions.
  • Practitioners have no guidance on when the resulting ability estimates and item parameters can be trusted.

Key ideas

  • Simulates 18,000 conditions across six LLM benchmarks to audit four estimator families: MML, MCMC, variational inference, and neural PSN.
  • Classical estimators become computationally infeasible at benchmark scale.
  • Scalable estimators produce unreliable item and ranking inferences when the model set is small or skewed — exactly the situation in frontier evaluation.
  • Provides concrete sample-size and diagnostic guidance for when IRT is safe to use.

Why it matters for evals

  • It is a timely, rigorous audit of a method the field adopted faster than it validated, and it arrives while government evaluators are actively building on IRT.
  • The actionable output is the diagnostics: run them before publishing an IRT-aggregated capability claim.
  • Two versions in under ten days signals active field interest.
  • Caveat: a simulation-based preprint; conclusions are conditioned on the simulated benchmark structures and estimator implementations chosen.

Comments

No comments yet.