Are LLM Benchmarks Already Contaminated? A Systematic Review
peer-reviewed review · source date 2026-07-01 · added 2026-08-07 15:52:43 · updated 2026-08-07 15:52:43 · Open original blog
1
Problems / challenges / motivations
- Contamination is the assumed explanation for suspicious benchmark gains, but the evidence was scattered across dozens of individual studies.
- Detection methods differ in what access they need (weights, logits, training data) and what kind of leakage they can see.
- Teams have no standard way to disclose what they did or did not check.
2
Key ideas
- A peer-reviewed systematic review of 55 contamination studies through late 2025.
- Proposes a four-tier taxonomy: Exact, Syntactic, Semantic, and Task-Level contamination.
- Compares five families of detection methods across access levels and training stages.
- Central finding: no detection method is reliable across all tiers, with observed score inflation in the rough range of 6–40%.
- Proposes a Contamination Transparency Card as a disclosure framework.
3
Why it matters for evals
- It is the most authoritative in-window synthesis of the problem underlying the month's validity reckoning, and unlike vendor blogs it is peer-reviewed.
- The Transparency Card is the practical takeaway: if you cannot prove a benchmark is clean, at least document what you checked and at which tier.
- Caveat: a review, not new measurement. The 6–40% band aggregates heterogeneous studies with different methods and should be read as indicative, not precise.
Comments
No comments yet.