AI & Agent Evaluation
2,151total visitsadmin

Are LLM Benchmarks Already Contaminated? A Systematic Review

peer-reviewed review · source date 2026-07-01 · added 2026-08-07 15:52:43 · updated 2026-08-07 15:52:43 · Open original blog

Problems / challenges / motivations

  • Contamination is the assumed explanation for suspicious benchmark gains, but the evidence was scattered across dozens of individual studies.
  • Detection methods differ in what access they need (weights, logits, training data) and what kind of leakage they can see.
  • Teams have no standard way to disclose what they did or did not check.

Key ideas

  • A peer-reviewed systematic review of 55 contamination studies through late 2025.
  • Proposes a four-tier taxonomy: Exact, Syntactic, Semantic, and Task-Level contamination.
  • Compares five families of detection methods across access levels and training stages.
  • Central finding: no detection method is reliable across all tiers, with observed score inflation in the rough range of 6–40%.
  • Proposes a Contamination Transparency Card as a disclosure framework.

Why it matters for evals

  • It is the most authoritative in-window synthesis of the problem underlying the month's validity reckoning, and unlike vendor blogs it is peer-reviewed.
  • The Transparency Card is the practical takeaway: if you cannot prove a benchmark is clean, at least document what you checked and at which tier.
  • Caveat: a review, not new measurement. The 6–40% band aggregates heterogeneous studies with different methods and should be read as indicative, not precise.

Comments

No comments yet.