CAISI Assessment of Z.ai's GLM-5.2
government evaluation · source date 2026-07-17 · added 2026-08-07 15:52:43 · updated 2026-08-07 15:52:43 · Open original blog
1
Problems / challenges / motivations
- Open-weight releases from outside the US arrive with vendor-reported numbers and no independent baseline.
- Aggregating many heterogeneous benchmarks into one capability claim is usually done informally.
- Contamination makes raw public-benchmark comparisons unreliable.
2
Key ideas
- A full US-government evaluation of the PRC open-weight GLM-5.2 across capabilities, cyber, safeguards, and security.
- Aggregation uses an Item-Response-Theory-style capability metric rather than averaging benchmark percentages.
- Findings: GLM-5.2 is roughly on par with GPT-5.2 overall and comparable to Opus 4.6 on cyber, with mixed safeguards results.
- The full report is published as a public PDF on NIST.gov.
3
Why it matters for evals
- This is a gold-standard example of contamination-resistant third-party capability measurement — the counterweight the field needs against vendor self-reports.
- The IRT aggregation is the transferable piece: it makes the weighting of items explicit and auditable instead of hiding it in a mean.
- Caveat: benchmark selection and the IRT aggregation still reflect CAISI's methodology choices, and this is a point-in-time assessment of one release.
Comments
No comments yet.