UK AISI + CAISI: Preliminary Assessment of Kimi K3's Cyber Capabilities
government evaluation · source date 2026-07-23 · added 2026-08-07 15:52:43 · updated 2026-08-07 15:52:43 · Open original blog
1
Problems / challenges / motivations
- Open-weight frontier releases cannot be recalled, so cyber risk assessment needs to happen before weights are public.
- Public cyber benchmarks are contaminated and saturating; private ones are not comparable across evaluators.
- No single government has full coverage of the frontier.
2
Key ideas
- A joint UK AISI + US CAISI pre-open-weight-release evaluation of Moonshot AI's Kimi K3 on public plus private benchmarks.
- Kimi K3 leads open-weight peers (32% versus GLM-5.2's 24%) but trails leading US closed models.
- On a 32-step attack path, K3 reaches step 17 versus 28.5 for the leading closed models — a metric that grades progress along a realistic kill chain rather than a single pass/fail.
- The assessment is mirrored on NIST.gov.
3
Why it matters for evals
- A bilateral government pre-release evaluation is rare, and it is the credible independent counterweight to Kimi K3's vendor-reported tables.
- The attack-path / TLO framing is a good example of partial-credit grading in a safety context: how far along the chain a model gets is more informative than whether it finished.
- Caveat: explicitly preliminary and cyber-scoped, published before Moonshot's final technical report. Treat it as an early independent signal, not a complete capability profile.
Comments
No comments yet.