GPT-5.6 System Card (Sol / Terra / Luna)
system card · source date 2026-07-09 · added 2026-08-07 15:52:43 · updated 2026-08-07 15:52:43 · Open original blog
1
Problems / challenges / motivations
- Standard safety benchmarks saturate, so a card built on them stops distinguishing models or catching regressions.
- Dangerous-capability thresholds (cyber, bio) need methodology that can rule things out, not just report a score.
- Smaller family members ship on the same weekend as the flagship and need their own risk assessment.
2
Key ideas
- Production Benchmarks with Challenging Prompts uses `not_unsafe` as the primary metric and is run without system-level safeguards, so the model itself is what is being measured.
- Coverage includes HealthBench and HealthBench Professional, chain-of-thought controllability, and the Preparedness evaluations.
- All three models are rated High in Cybersecurity and Biological/Chemical — the first time smaller family members received a High designation — and below High in AI self-improvement.
- Cyber Critical is ruled out via high test-time-compute exploitation tests and staged verifier oracles.
3
Why it matters for evals
- It is a canonical example of production-representative safety benchmarking replacing saturated academic sets.
- The staged-verifier-oracle and CTF-proxy methodology for ruling out a capability threshold is now referenced across the field; ruling out is harder to do well than measuring.
- Caveat: capability and Preparedness numbers are self-reported from OpenAI's harness. Pair with independent AISI/CAISI evaluations and account for eval-awareness confounds.
Comments
No comments yet.