AI & Agent Evaluation
2,151total visitsadmin

Gemini 3.6 Flash — Model Card and Launch

system card · source date 2026-07-21 · added 2026-08-07 15:52:43 · updated 2026-08-07 15:52:43 · Open original blog

Problems / challenges / motivations

  • Fast, cheap models are where most production agent traffic actually runs, but they get thinner eval treatment than flagships.
  • Agentic and coding capability is scaffold-dependent, so a number without its harness is close to meaningless.
  • Safety-framework reporting needs to happen for every release, not only the largest one.

Key ideas

  • The card documents evaluation across reasoning, coding, agentic, multimodal, and long-context benchmarks: DeepSWE, MLE-Bench, Terminal-Bench 2.1, OSWorld-Verified, GDPval-AA v2, and GDM-MRCR v2.
  • It includes automated safety, tone, and refusal evaluations plus a Frontier Safety Framework assessment.
  • Results are reported as deltas versus Gemini 3.5 Flash, with an emphasis on token efficiency rather than raw score.

Why it matters for evals

  • Token efficiency reported alongside accuracy is the right framing for the Flash class, and it is a reminder that cost-normalized comparison belongs in the card rather than in third-party analysis afterward.
  • It gives Google/DeepMind a fresh in-window frontier card with detailed agentic tables, and it was immediately re-analyzed across the ecosystem.
  • Caveat: benchmark numbers are vendor-reported from Google's own harness and scaffolds. Leaderboard-style figures (Terminal-Bench 2.1, GDPval-AA) are harness-dependent and should not be compared cross-lab without matching setups.

Comments

No comments yet.