AI & Agent Evaluation
2,151total visitsadmin

Kimi K3 — Launch and Evaluation Disclosures

launch coverage · source date 2026-07-16 · added 2026-08-07 15:52:43 · updated 2026-08-07 15:52:43 · Open original blog

Problems / challenges / motivations

  • The largest open-weight releases now ship with vendor eval tables well before any technical report or independent check exists.
  • Practitioners have to decide whether to trust those tables in the gap between launch and verification.
  • Open-weight releases are irreversible, which raises the stakes on that gap.

Key ideas

  • Kimi K3 is a 2.8T-parameter model (16 of 896 experts active) with a 1M-token context, released as open weights.
  • Vendor eval tables report GPQA-Diamond 93.5, Terminal-Bench 2.1 88.3, BrowseComp 91.2, MCP-Atlas 84.2, and SWE-Marathon 42.0.
  • Moonshot claims top open-weight standing, behind only GPT-5.6 Sol and Claude Fable 5.
  • Full weights and the technical report were promised for 2026-07-27, after the launch numbers circulated.

Why it matters for evals

  • It is a live case study in eval-disclosure practice: the numbers arrived first, the report second, and independent verification third.
  • Read it against the UK AISI + CAISI cyber assessment of the same model, which is the credible independent signal and lands more conservatively.
  • Caveat: numbers are vendor-reported and pre-technical-report. This link is credible practitioner coverage (The Batch), not the primary Moonshot report.

Comments

No comments yet.