AI & Agent Evaluation
2,151total visitsadmin

CAISI Assessment of Z.ai's GLM-5.2

government evaluation · source date 2026-07-17 · added 2026-08-07 15:52:43 · updated 2026-08-07 15:52:43 · Open original blog

Problems / challenges / motivations

  • Open-weight releases from outside the US arrive with vendor-reported numbers and no independent baseline.
  • Aggregating many heterogeneous benchmarks into one capability claim is usually done informally.
  • Contamination makes raw public-benchmark comparisons unreliable.

Key ideas

  • A full US-government evaluation of the PRC open-weight GLM-5.2 across capabilities, cyber, safeguards, and security.
  • Aggregation uses an Item-Response-Theory-style capability metric rather than averaging benchmark percentages.
  • Findings: GLM-5.2 is roughly on par with GPT-5.2 overall and comparable to Opus 4.6 on cyber, with mixed safeguards results.
  • The full report is published as a public PDF on NIST.gov.

Why it matters for evals

  • This is a gold-standard example of contamination-resistant third-party capability measurement — the counterweight the field needs against vendor self-reports.
  • The IRT aggregation is the transferable piece: it makes the weighting of items explicit and auditable instead of hiding it in a mean.
  • Caveat: benchmark selection and the IRT aggregation still reflect CAISI's methodology choices, and this is a point-in-time assessment of one release.

Comments

No comments yet.