AI & Agent Evaluation
2,151total visitsadmin

Claude Opus 5 System Card

system card · source date 2026-07-24 · added 2026-08-07 15:52:43 · updated 2026-08-07 15:52:43 · Open original blog

Problems / challenges / motivations

  • A frontier model release needs a single artifact that shows what was tested before deployment, not just headline capability scores.
  • Cyber and bio risk assessments are hard to make credible when the lab designs, runs, and reports its own evaluations.
  • Alignment claims need something more structured than spot checks on refusals.

Key ideas

  • The card is a full pre-deployment eval dossier: capability, cyber (ExploitBench, OSS-Fuzz, Firefox 147, plus new CyScenarioBench and ExploitGym), safeguards and harmlessness, an automated behavioral audit, and model welfare.
  • Anthropic brought in an external UK AISI cyber range rather than relying only on in-house cyber tests.
  • Reported results include an overall misaligned-behavior score of 2.3 — described as Anthropic's "most aligned model to date" — and a CB-1 bio-risk designation.
  • The automated behavioral audit is notable as a scalable substitute for hand-run red-team sessions.

Why it matters for evals

  • This is the reference template for how a complete eval suite gets structured and reported at the frontier, and it was widely cited across launch coverage.
  • The third-party cyber-range component is the part worth copying: it is the difference between a self-graded exam and one with an external proctor.
  • Caveat: capability and safety numbers are vendor self-reported from Anthropic's own harness. Pair with independent government evaluations (UK AISI, CAISI) and treat eval awareness and contamination as live confounds when reading any single number.

Comments

No comments yet.