Google DeepMind introduced Gemini 4 Argon on September 30, 2026, calling it its new frontier model, "built for complex workflows across coding, enterprise knowledge work, and cybersecurity defense." It is rolling out to trusted testers and cyber defenders through Google's Fairwind Program. In Google's own 19-benchmark table, Argon has the best or tied-best score on 14 rows (outright best on 13). It is compared with GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5, all publicly available.
Benchmark table from Google DeepMind's Gemini 4 Argon announcement.
Already working inside Google
Google says thousands of Googlers already use Argon for specialized coding, research and writing. Examples:
- Quantum algorithms: Argon optimized the spacetime resources (qubits x gates) of bottleneck subroutines, beating the published baseline by 40% in one case.
- Memory: Argon agents analyzed fleet-wide profiling telemetry and applied optimizations across Google's data centers, freeing over 300 TiB (an estimated 500 TiB to 1 PiB in total).
- Rust migrations: agents are migrating C/C++ to Rust, from tens of thousands of lines in libraries such as re2 and libgav1 up to 800K+ lines for the Fuchsia Zircon kernel. These rewrites are under auditing and review before production.
- libgav1: agents replaced 32K lines of SIMD code in an existing Rust port, producing a memory-safe video decoder that runs 2.7x faster than the port with identical output.
Where it leads
Google headlines state of the art on DeepSWE v1.1 (77.9%), leadership on the Vals Index, Vals Finance Agent v2 and Harvey's Legal Agent Benchmark, the #1 spot on Zapier's AutomationBench (51.3%) and state of the art on LVBench (91.7%). Its output limit rises to 1 million tokens from 64K.
The widest margins are in knowledge work and long context. On AutomationBench, Astra scores 41.4%, Fable 5.1 31.4% and Opus 5.5 42.5%. On Harvey's benchmark Argon scores 19.6% against 5.4%, 6.7% and 3.8%. It leads Vals Finance Agent v2 at 65.4% (next best 58.9%) and the Vals Index at 68.9% (Opus 67.0%). On GraphWalks at 256k to 1M tokens it reaches 84.2 F1 against 71.8, 65.0 and 66.8.
It also leads LABBench 2 (88.8%), RiemannBench (76.0%), Chartography (71.6%) and Vibe Code Bench (91.9%). On CWE-bench v1, Argon ties Astra for first at 68.0%, ahead of Fable 5.1 (58.0%) and Opus 5.5 (67.0%).
Per rival, counting only rows where the rival has a score:
- GPT-6 Astra: 14 wins, 4 losses, 1 tie over 19 rows.
- Claude Fable 5.1: 15 wins, 2 losses over 17 rows.
- Claude Opus 5.5: 14 wins, 4 losses over 18 rows.
Where it trails
Agentic coding is the main weak spot, with OSWorld-2.0 also trailing Astra. Argon is 10.5 points behind Astra on FrontierSWE v2 and 9.0 behind Opus 5.5 on Terminal-bench 4.0, and Fable 5.1 also edges it on both. The five rows where another model leads:
| Benchmark | Gemini 4 Argon | Leader | Leader's score |
|---|---|---|---|
| FrontierSWE v2 | 55.0% | GPT-6 Astra | 65.5% |
| Terminal-bench 4.0 | 57.4% | Claude Opus 5.5 | 66.4% |
| PostTrainBench | 45.3% | Claude Opus 5.5 | 49.3% |
| Terminal-Bench Science 0.1 | 57.6% | GPT-6 Astra | 68.1% |
| OSWorld-2.0 (offline subset, partial score) | 69.2% | GPT-6 Astra | 72.6% |
Fable 5.1 and Opus 5.5 have no reported OSWorld-2.0 score. Bloomberg reported that some Google employees say Gemini 4 does well on benchmarks but struggles with some real-world coding tasks; Google disputes that.
Cyber defense and safeguards
Google trained Argon to find, validate and patch critical vulnerabilities autonomously. Trusted defenders and Google's internal teams get it without cyber guardrails. Wiz, through its Scan for Good initiative, used Argon to uncover a critical vulnerability exposing sensitive personal information in healthcare software used by hospitals worldwide, a risk previous frontier models had missed. Argon also outperforms Gemini 3.8 Flash Cyber on Google's internal vulnerability benchmark, which spans 20 programming languages, and on Wiz's black-box penetration testing benchmark.
Google describes four safeguard areas:
- Misuse: refusal of harmful cyber and CBRN requests plus monitoring of the model's internal activations, tested by internal and external red teams.
- Prompt injection: Argon is Google's most resilient model yet against indirect prompt injection and leads Gray Swan's IPI benchmark.
- Misalignment: monitoring of chain-of-thought and actions can stop execution. A similar system monitored training runs, with findings kept out of training so they do not teach the model to evade monitoring. Google urges the industry to preserve reasoning transparency.
- Hardening: sandboxes are isolated and sealed before high-risk training or evaluations.
Access
Fairwind launched on September 2 alongside Gemini 3.8 Flash Cyber and the CodeMender harness. The program is open to trusted government authorities, critical-infrastructure operators and software maintainers. Anthropic similarly limits Claude Mythos 5.1 to vetted organizations, as covered in this blog's Fable 5.1 and Mythos 5.1 post; here the gated model is a general frontier model rather than a cyber variant.
Introductory pricing is $2 per million input tokens and $10 per million output tokens, with cached input 95% off. Google says it is engaged in the U.S. government's voluntary process for pre-release model access while it gradually expands access. Wider availability starts with paid API customers and Google AI Ultra subscribers, with no general-availability date given.
The New Stack's headline puts it this way: "Gemini 4 Argon is here, it's great, and you can't have it yet."