Google Introduces Gemini 4 Argon, Beats GPT 6 Astra And Opus 5.5 On Many Benchmarks

After more than half a year in the wilderness, Google is back in the AI game.

Google has announced Gemini 4 Argon, its new frontier model, and the company’s own numbers have it ahead of rivals on a majority of the benchmarks it chose to publish. Argon is rolling out first to a set of trusted cyber defenders through Google’s Fairwind Program, with a wider release to developers, enterprises and consumers to follow once feedback from early testers has been folded into the model’s safeguards.

In a blog post, Google DeepMind’s Koray Kavukcuoglu described Argon as a model built to sustain deep reasoning across long-horizon workflows, with strengths in real-world software engineering, enterprise knowledge work such as legal and finance, and cybersecurity defense. In Google’s comparison table, Argon is set against OpenAI’s GPT-6 Astra, Anthropic’s Claude Fable 5.1 and Claude Opus 5.5.

What Google Is Announcing

Argon is not yet broadly available. Google says it is taking a phased approach, and is currently taking part in the U.S. government’s voluntary process for pre-release model access while it gradually expands availability. The first users are trusted cyber defenders, who will get a version of the model without cyber guardrails so they can use its full defensive capabilities. The same unrestricted version is being used by Google’s internal teams.

After that, Google plans to open Argon up to paid API customers and Google AI Ultra subscribers first, and then to a broader set of developers, enterprises and consumers.

The pricing is aggressive. Argon will launch at an introductory $2 per million input tokens and $10 per million output tokens, with cached input tokens discounted 95% from the input price. That undercuts Anthropic’s Claude Opus 5.5, which is listed at $4 and $20 per million tokens, and sits well below the $10 and $50 charged for both GPT-6 Astra and Claude Fable 5.1. The “introductory” label is worth noting, since Google hasn’t said what happens to pricing after the launch period.

Another headline change is the output limit. Argon supports up to 1 million output tokens in a single response, up from 64K previously. Google says this gives the model room to think for hundreds of thousands of tokens in a single trajectory and tackle hard problems in one pass.

Gemini 4 Argon Benchmarks

Google’s table covers 19 benchmarks across knowledge work, agentic coding, ML engineering, science and math, long context, computer use, multimodal understanding and cybersecurity. Argon is first or tied for first on 14 of them, and trails a rival on the other five. Here’s how it breaks down.

Knowledge work is the clearest win

This is where Argon looks most dominant. It leads all four knowledge-work benchmarks, and in three of them the margin is large. On AutomationBench, Zapier’s test of end-to-end business task execution, Argon scores 51.3%, roughly nine points clear of the next-best model, Opus 5.5 at 42.5%, and nearly 20 points ahead of Claude Fable 5.1’s 31.4%. On Vals Finance Agent v2, which measures multi-step financial research, Argon’s 65.4% is almost seven points ahead of Fable 5.1, the closest competitor.

The Vals Index, which weighs finance, coding, legal and tax work by each sector’s contribution to U.S. GDP, is much tighter. Argon’s 68.9% is under two points ahead of Opus 5.5’s 67.0%, and Google calls Argon the leading model on it. Harvey’s Legal Agent Benchmark produces the most eye-catching gap in the table, with Argon at 19.6% against 6.7% for Fable 5.1 and just 5.4% for Astra. The absolute scores are low across the board, which suggests the benchmark is still very hard for every model, but Argon is roughly three times better than the best of its rivals here.

Coding: a new state of the art on DeepSWE, but not a sweep

Google says Argon sets a new state of the art on DeepSWE v1.1, which measures long-horizon real-world software engineering, with 77.9%. Astra (74.1%) and Opus 5.5 (74.2%) are about three and a half points behind, while Fable 5.1 sits at 67.4%. Argon also tops Vibe Code Bench at 91.9%, though the field is bunched within about two points there.

But coding is also where Argon’s weaknesses show up. On FrontierSWE v2, Argon scores 55.0%, which is the lowest of the four models and ten and a half points behind Astra’s 65.5%. On Terminal-Bench 4.0, it manages 57.4%, behind Astra, Fable 5.1 and, by a wide margin, Opus 5.5 at 66.4%. Opus 5.5 also leads on PostTrainBench, the table’s lone ML engineering test, at 49.3% against Argon’s 45.3%. Anthropic’s model remains the one to beat for terminal-style agentic work, at least on these numbers.

Science and math: strong on LABBench and RiemannBench, weaker on terminal science

Argon leads on LABBench 2 (88.8%) with a particularly wide gap over Opus 5.5 (73.1%) and Fable 5.1 (68.6%), and on RiemannBench (76.0%), where it is four points clear of Astra. The exception is Terminal-Bench Science 0.1, where Astra takes the top spot at 68.1% against Argon’s 57.6%, with Opus 5.5 also ahead at 63.3%. That pattern, with Argon weaker on terminal-driven agentic tasks across both the coding and science categories, is the most consistent theme in where it loses.

Long context is a standout

The long-context results may be the most striking in the table. On GraphWalks up to 128K tokens, Argon scores 99.7%, which is close to a ceiling, but the real separation comes at longer lengths. At 256K to 1M tokens, Argon scores 84.2% versus 71.8% for Astra, 66.8% for Opus 5.5 and 65.0% for Fable 5.1. A gap of more than 12 points over the nearest competitor at the longest context lengths suggests Google’s model degrades far less as inputs grow, which matters for the long-horizon workflows Google is pitching.

Computer use is mixed

Argon leads on Agent’s Last Exam with a 39.5% pass rate, ahead of Opus 5.5 (38.2%) and Astra (34.2%), though no score is listed for Fable 5.1. On OSWorld-2.0’s offline subset, which uses a partial-score metric, Astra edges ahead at 72.6% to Argon’s 69.2%. Neither Fable 5.1 nor Opus 5.5 has a score listed there, so the comparison is a two-horse race.

Multimodal: a big gap on Chartography and LVBench

Google highlights Argon’s strength on tasks that need visual understanding, and the table backs that up. On LVBench, which tests long video understanding, Argon’s 91.7% is more than four points ahead of Astra (87.5%) and eight clear of Opus 5.5 (83.7%). On Chartography, Argon (71.6%) is just ahead of Astra (71.0%), but both are well clear of Opus 5.5 (66.3%) and Fable 5.1 (46.2%).

Cybersecurity: a tie at the top

On CWE-bench v1, which evaluates a model’s ability to remediate security vulnerabilities, Argon ties Astra for first place at 68.0%. Opus 5.5 is only a point behind at 67.0%, while Fable 5.1 trails at 58.0%. Google says the result builds on the frontier performance of Gemini 3.8 Flash Cyber on the earlier CWE-bench v0.

Beyond the table, Google says Argon uncovered a wide range of exposures on its internal vulnerability benchmark across codebases in 20 programming languages. It also says the model outperforms 3.8 Flash Cyber on Wiz’s internal black-box penetration testing benchmark. Wiz is already using Argon through its Scan for Good initiative, and says the model found a critical vulnerability exposing sensitive personal information in healthcare software used by hospitals around the world, one that earlier frontier models had missed.

Early Internal Results

Google also shared how Argon is being used inside the company, with thousands of Googlers using it for specialized coding, deeper research and writing:

  • Quantum algorithm optimization: Argon helped researchers cut the spacetime resources (qubits × gates) of subroutines that bottleneck important applications, beating a published baseline by 40% in a matter of minutes.
  • Memory efficiency: A team of Argon agents analyzed fleet-wide profiling telemetry and applied optimizations that Google says will free up over 300 TiB of memory across its data centers once rolled out, with total savings estimated at 500 TiB to 1 PiB.
  • Codebase migrations: Argon agents are migrating C and C++ code to Rust, from tens of thousands of lines in libraries like re2 and libgav1 up to more than 800,000 lines for the Fuchsia Zircon kernel. The work is undergoing automated and manual audits, emulation testing and review before reaching production.
  • libgav1: Starting from an existing Rust port, Argon agents replaced 32K lines of SIMD code with safe Rust the compiler could vectorize on its own, producing a memory-safe video decoder that runs 2.7x faster than the Rust port with identical output.

Safeguards Before Wider Release

Google says it is strengthening four areas before a broad launch. For misuse, it is improving techniques that monitor the model’s internal activations to catch cyber and CBRN abuse, with red teaming from internal and external groups. On prompt injection, Google says Argon is its most resilient model so far, leading Gray Swan’s Indirect Prompt Injection benchmark.

For misalignment, Google is deploying monitors that watch Argon’s chain of thought and actions and halt execution when needed. It says it is deliberately avoiding feeding monitoring findings back into training so as not to teach the model’s reasoning to evade oversight, and urged the wider industry to preserve reasoning transparency. Finally, Google is hardening its sandboxed environments by isolating and sealing them before high-risk training or evaluations begin.

Posted in AI