AI & Agent Evaluation
2,151total visitsadmin

UK AISI + CAISI: Preliminary Assessment of Kimi K3's Cyber Capabilities

government evaluation · source date 2026-07-23 · added 2026-08-07 15:52:43 · updated 2026-08-07 15:52:43 · Open original blog

Problems / challenges / motivations

  • Open-weight frontier releases cannot be recalled, so cyber risk assessment needs to happen before weights are public.
  • Public cyber benchmarks are contaminated and saturating; private ones are not comparable across evaluators.
  • No single government has full coverage of the frontier.

Key ideas

  • A joint UK AISI + US CAISI pre-open-weight-release evaluation of Moonshot AI's Kimi K3 on public plus private benchmarks.
  • Kimi K3 leads open-weight peers (32% versus GLM-5.2's 24%) but trails leading US closed models.
  • On a 32-step attack path, K3 reaches step 17 versus 28.5 for the leading closed models — a metric that grades progress along a realistic kill chain rather than a single pass/fail.
  • The assessment is mirrored on NIST.gov.

Why it matters for evals

  • A bilateral government pre-release evaluation is rare, and it is the credible independent counterweight to Kimi K3's vendor-reported tables.
  • The attack-path / TLO framing is a good example of partial-credit grading in a safety context: how far along the chain a model gets is more informative than whether it finished.
  • Caveat: explicitly preliminary and cyber-scoped, published before Moonshot's final technical report. Treat it as an early independent signal, not a complete capability profile.

Comments

No comments yet.