Gemini 3.6 Flash — Model Card and Launch
system card · source date 2026-07-21 · added 2026-08-07 15:52:43 · updated 2026-08-07 15:52:43 · Open original blog
1
Problems / challenges / motivations
- Fast, cheap models are where most production agent traffic actually runs, but they get thinner eval treatment than flagships.
- Agentic and coding capability is scaffold-dependent, so a number without its harness is close to meaningless.
- Safety-framework reporting needs to happen for every release, not only the largest one.
2
Key ideas
- The card documents evaluation across reasoning, coding, agentic, multimodal, and long-context benchmarks: DeepSWE, MLE-Bench, Terminal-Bench 2.1, OSWorld-Verified, GDPval-AA v2, and GDM-MRCR v2.
- It includes automated safety, tone, and refusal evaluations plus a Frontier Safety Framework assessment.
- Results are reported as deltas versus Gemini 3.5 Flash, with an emphasis on token efficiency rather than raw score.
3
Why it matters for evals
- Token efficiency reported alongside accuracy is the right framing for the Flash class, and it is a reminder that cost-normalized comparison belongs in the card rather than in third-party analysis afterward.
- It gives Google/DeepMind a fresh in-window frontier card with detailed agentic tables, and it was immediately re-analyzed across the ecosystem.
- Caveat: benchmark numbers are vendor-reported from Google's own harness and scaffolds. Leaderboard-style figures (Terminal-Bench 2.1, GDPval-AA) are harness-dependent and should not be compared cross-lab without matching setups.
Comments
No comments yet.