2026-09-02

Gemini 3.8 Flash and Flash Cyber: the Third Flash Release in Six Weeks, and the Fourth Lab With a Cyber Model

AISecurity🌍 North America

Google shipped Gemini 3.8 Flash and Gemini 3.8 Flash Cyber on September 2, 2026, three weeks after Gemini 3.7 Flash — by Google's own count, its third Flash release in six weeks. Both variants share the same underlying model; what differs is which capabilities are switched on and who can reach them.

Gemini 3.8 Flash: eight wins out of fourteen, but not the hardest ones

Google published a 14-benchmark comparison table against Gemini 3.7 Flash, Claude Opus 5, Claude Sonnet 5, GPT-5.6 Sol, and GPT-5.6 Terra. Gemini 3.8 Flash leads outright on eight: Vals Finance Agent V2 (61.4%), Harvey's Legal Agent Benchmark (10.0% — the whole field clusters in single digits here), Terminal-bench 2.1 (89.4%), CharXiv Reasoning (86.2%), LVBench (87.8% agentic / 87.1% static), HLE-Verified (54.9%), the harder split of BioMysteryBench (56.5%), and LABBench2 (86.2%).

Claude Opus 5 — a frontier-tier model, not a Flash-class rival — takes the other five: DeepSWE v1.1 long-horizon software engineering (74.0% against 3.8's 73.7%), GDPVal-AA v2 knowledge-work Elo (1824 against 1545), Terminal-bench 4.0 general agent capabilities (51.8% against just 19.1% — the widest gap anywhere in the table), OSWorld-2.0 agentic computer use (75.4% against 59.0%), and the easier split of BioMysteryBench (90.1%). GPT-5.6 Sol takes the remaining row, GDP.PDF document comprehension (40.0%).

The shape of that split matters more than the 8-to-5 scoreline. Every one of Opus 5's wins is on a benchmark measuring sustained, general-purpose agentic capability rather than a specific professional task — and on the one built explicitly to be hard (Terminal-bench 4.0), the gap is enormous. Gemini 3.8 Flash's real claim is winning the majority of individual rows at Flash-tier pricing, not closing that gap. Google's own explanation of the mechanism fits: the model "works harder" on complex tasks — more reasoning steps, more iterative tool calls — which can mean more tokens burned at higher effort levels than 3.7 Flash used on the same task. For workloads where compute is the binding constraint, Google's answer is to drop the effort level rather than the model; 3.7 Flash also stays fully supported for anyone who wants the old cost profile.

The pricing itself confirms what we flagged when 3.7 Flash launched: the $0.75/$3.75 rate is introductory for both 3.7 and 3.8 Flash, expiring December 31, 2026, after which both revert together to $1.50/$7.50 on January 1, 2027. It's not a permanent cut, and it was never exclusive to the new model.

Flash Cyber: the fourth lab to ship this exact shape

Three frontier labs shipped a dedicated cybersecurity model within one summer — OpenAI's GPT-5.5-Cyber in May, Sakana's Fugu-Cyber and Google's own Gemini 3.5 Flash Cyber both on the same day in July. Gemini 3.8 Flash Cyber is Google's second entry in that lineage, and the access mechanism has changed: 3.5 Flash Cyber went out through Google's CodeMender agent to governments and trusted partners; 3.8 Flash Cyber ships through a new Fairwind Program, opened to trusted government authorities, critical infrastructure operators, and software maintainers specifically.

On CyberGym, the standard industry vulnerability-discovery benchmark, Google says 3.8 Flash Cyber surpasses both 3.5 Flash Cyber and "significantly larger frontier models" — a qualitative claim rather than a published number in the announcement itself. On an internal 20-programming-language vulnerability benchmark built to better match real-world codebases than CyberGym's C/C++ focus, Google reports a success rate over 70%, calling it "an impressive leap" over its previous models. On CWE-Bench, a third-party patching benchmark run by Collinear, Google frames 3.8 Flash Cyber as sitting on the Pareto frontier: a 47.2% pass@1 against 47.8% for a leading frontier model, at what Google describes as significantly lower cost.

The more specific design choice is where Google put its effort: patching over exploitation. The company says it invested in vulnerability fixing from the start and prioritized that over offensive capability — the same split Anthropic drew for Claude Fable 5.1 and Mythos 5.1 one day earlier, which allows Fable 5.1 to discover vulnerabilities but routes exploit development to a more restricted model. Two labs landing on identical framing — discovery yes, exploitation no — within 24 hours of each other suggests this is becoming the default posture for cyber-capable models generally, not a one-off decision at either company.

Third-party validation, with real numbers attached

Unlike the benchmark claims above, the case studies Google cites carry specific figures. Chrome's Security team reports 3.8 Flash Cyber produced 2.6 times more correct patches for Chrome vulnerabilities than "the best commercial models that are much larger." Wiz reports 7.5–9.7% higher recall on its internal penetration-testing benchmark at 2.3–5.2 times lower cost than other leading frontier models. Google's own Cloud Vulnerability Research team says it used the model to find a critical foundational vulnerability in under two hours — research Google says usually takes months. These are still vendor-selected examples rather than independent audits, but they're at least attached to named teams and concrete multipliers rather than a chart alone.

Safety framing

Gemini 3.8 Flash ships with safeguards against CBRN and cyber-offense misuse under Google's Frontier Safety Framework. The Cyber variant runs a deliberately more permissive set of mitigations, which is exactly why it's gated behind the Fairwind Program rather than shipped as a public API model. Separately, Google says both 3.8 models made "a significant leap" in prompt-injection robustness as measured by Gray Swan, though again without a specific score quoted in the announcement text.

Availability

3.8 Flash is live today for developers in Google Antigravity, the Gemini API via AI Studio and Android Studio, and Stitch; for enterprises through Gemini Enterprise; and for consumers on Google AI Pro and Ultra across the Gemini app, AI Mode in Search, and Gemini in Sheets. Flash Cyber access runs through an application to the Fairwind Program.

What to expect next

  • Watch for CyberGym's actual pass@1 number. Google quotes a qualitative lead over 3.5 Flash Cyber's 83.2% and over larger frontier models, but the announcement doesn't publish the figure directly — worth checking back once it surfaces on the public leaderboard or in independent coverage.
  • Watch whether "discover but don't exploit" becomes the industry standard split. Anthropic and Google landing on the same design boundary within a day of each other could be convergent engineering or could be the two companies watching the same regulatory signals — worth checking whether OpenAI's and Sakana's next cyber-model updates draw the same line.
  • Watch the Fairwind Program's actual admission criteria. "Trusted government authorities, critical infrastructure operators, and software maintainers" is broader than CodeMender's "governments and trusted partners" — whether that's real expansion of access or just clearer language will show up in who actually gets in.