2026-08-13

Gemini 3.7 Flash Wins Nine of Twenty Benchmarks and Still Doesn't Lead the Composite Score

AIBenchmarks🌍 North America

Google shipped Gemini 3.7 Flash today, three weeks after Gemini 3.6 Flash — the same tight release cadence we've watched from xAI and DeepSeek this week. Google's framing is "our most intelligent workhorse model yet for coding and agents," which is the kind of line that means nothing until you check it against the comparison table Google published alongside it, against Gemini 3.6 Flash, Claude Sonnet 5, GPT-5.6 Terra, and Muse Spark 1.2.

The table says something more specific than the headline

Across 20 scored benchmarks, Gemini 3.7 Flash leads outright on nine — more than any of the four rivals it's compared against: FrontierCode 1.1 Main (43.6% vs 3.6's 34.4%), Code Arena web-dev Elo (1588), AutomationBench private set (30.4%), Harvey LAB-AA legal workflows (90.7%), GDP.pdf document comprehension (34.0%), LVBench long-video (85.4%), GDM-MRCR v2 long-context (97.0%), HLE-Verified (53.6%), and LABBench2 biology research tasks (82.1%). That's a genuinely broad win column, not a cherry-picked pair of categories.

And yet Gemini 3.7 Flash doesn't lead the Artificial Analysis Intelligence Index — the single composite score at the top of Google's own table. It scores 56, tied for last among the five models shown; GPT-5.6 Terra and Muse Spark 1.2 both score 57. That's the inverse of the pattern we flagged in Grok 4.6's launch, where the headlined composite metric flattered the model more than the underlying table did. Here the composite undersells it: a model that wins nearly half the individual categories in its own comparison sheet posts the lowest aggregate score of the five. Composite indices average across everything a model is bad at along with everything it's good at, and Gemini 3.7 Flash's real story is concentration, not breadth — strong in document comprehension, long-context retrieval, and specific agentic workflows, weaker in exactly the cluster where it loses.

Where it loses, and it's a consistent cluster

GPT-5.6 Terra leads six benchmarks, and they're not scattered — they cluster tightly around long-horizon coding-agent work: DeepSWE v1.1 (69.6% vs Gemini 3.7's 65.3%), Terminal-bench 2.1 (87.4% vs 85.8%), Terminal-bench 3.0 (20.8% vs 14.9%), OSWorld-2.0 agentic computer use (50.2% vs 47.9%), CharXiv Reasoning without tools, and BioMysteryBench's harder tier. That's the same cluster that's shown up as the hard-to-compress capability all month — sustained, tool-using, long-horizon execution is where the frontier keeps separating from the workhorse tier, regardless of which lab's workhorse model it is. Claude Sonnet 5 takes two categories (Agent's Last Exam, BioMysteryBench's easier tier), and — genuinely notable — Gemini 3.6 Flash beats its own successor on one narrow metric, CharXiv Reasoning with tools (89.4% vs 3.7's 88.7%), a reminder that "newer" isn't a strict upgrade on every axis even within the same family.

The "half price" claim needs one clarification

Google's text says 3.7 Flash launches "at an introductory price of half the original 3.6 Flash cost per million tokens" — $0.75/$3.75 versus what the post implies was 3.6's standard $1.50/$7.50 (a figure that matches what we recorded when 3.6 Flash launched). What the comparison table actually shows is that Gemini 3.6 Flash is priced identically to 3.7 Flash right now — both marked with the same asterisk: this is an introductory rate for both models, expiring December 31, 2026, after which both revert to $1.50/$7.50. So the "half price" framing is accurate against 3.6's original, pre-promotion rate, but the practical effect today is that Google quietly cut 3.6 Flash's price to match 3.7's launch price too, rather than only pricing the new model aggressively. Existing 3.6 Flash customers get the same temporary discount as new 3.7 adopters — worth knowing before assuming the price cut is exclusive to the new model.

The product ecosystem around the benchmark table

A few details worth noting outside the scored comparisons: Google demoed 3.7 Flash paired with Nano Banana to generate a playable 3D game in real time from a text prompt, and with Gemini Omni orchestrating sub-agents for one-shot interactive landing pages — both existing Google products getting folded into 3.7 Flash's launch materials rather than benchmarked claims. A robotics demo used "a 3 agent graph loop" to help a robot learn faster, the same multi-agent-decomposition pattern we've now seen in AMIE's Talker/Planner/Perception split applied to an entirely different domain — one more data point that splitting an agent into specialized sub-agents is becoming a default architecture, not a one-off research choice.

More consequentially: Gemini Spark — the always-on personal agent Google launched at I/O — now runs on 3.7 Flash starting today. That's the same product category Grok Bot entered two days ago: a persistent background agent taking action on a user's behalf, now getting a quiet model swap rather than a headline launch of its own. Worth watching both companies' background-agent products converge on the same upgrade cadence — model improvements arriving as invisible infrastructure updates to a product a user already has open, rather than as something they have to opt into.

The safety section, once more, is adjectives

Consistent with every model launch this month, the capability claims come with exact numbers and the safety claims don't: "updated safeguards against misuse in the domains of Chemical, Biological, Radiological, and Nuclear (CBRN) and cyber offense," no figures, no named evaluation, a pointer to the model card rather than data in the post itself. Same asymmetry as Grok 4.6's launch, DeepSeek's V4-Pro, and effectively every other frontier release covered this month — worth naming as the pattern it now clearly is rather than a one-off omission.

What to expect next

  • Watch for independent replication of the nine category wins, especially Harvey LAB-AA and LABBench2 — domain-specific benchmarks in legal and biology work are exactly the kind of claim worth checking against a source outside the vendor before treating as settled.
  • Watch what happens to 3.6 Flash pricing on January 1, 2027. Both models reverting together to $1.50/$7.50 is a real, dated commitment — whether that date holds, or gets extended the way open-weights promises have been slipping this month, is worth checking back on.
  • Watch the coding-agent cluster specifically. GPT-5.6 Terra's lead on DeepSWE, Terminal-bench, and OSWorld is consistent enough across labs this month that it looks less like one model's edge and more like a genuinely harder capability frontier than the rest of the table.

References: Google — Introducing Gemini 3.7 Flash · Google — Gemini 3.6 Flash launch, for pricing baseline · related coverage: Grok 4.6 launch · DeepSeek-V4-Pro · Grok Bot · AMIE (Video) · Frontier Arcade: trends & predictions