Gemini 4 Argon Can Generate 1 M Output Tokens, Compared To 128k For Most Other Frontier Labs

Google’s new frontier model has a headline feature that has little to do with benchmark scores: it can write up to a million tokens in a single response. That is nearly eight times the 128K output ceiling listed for the flagship models from OpenAI and Anthropic.

Google DeepMind announced Gemini 4 Argon today, raising the model’s maximum output from the 64K limit of earlier Gemini models to 1 million tokens. Google says the extra room lets Argon reason for hundreds of thousands of tokens within one trajectory and work through hard problems in a single pass, instead of being stopped and re-prompted.

How Argon compares

Here is how the maximum output per response stacks up against the models Google positions Argon against, along with OpenAI’s newer GPT-6.1 Sol and Anthropic’s Claude Sonnet 5.5:

ModelMax output per response
Gemini 4 Argon1,000K
GPT-6.1 Sol128K
Claude Opus 5.5128K
Claude Sonnet 5.5128K
Claude Fable 5.1128K

At 1,000K against 128K, Argon’s ceiling is roughly 7.8 times higher. GPT-6 Astra, which Google also compares Argon with, is listed at the same 128K.

There is one nuance on the Anthropic side. Anthropic’s documentation says Sonnet 5 can produce up to 300K output tokens on its Message Batches API behind a beta header. That is still well short of 1M, and the standard limit across Anthropic’s current lineup remains 128K.

Context and output are not the same thing

The two numbers are easy to confuse. The context window is how much a model can take in: the documents, code, conversation history and tool results it can read at once. The output limit is how much it can generate in a single run, and that includes the tokens it spends reasoning before it answers.

Most frontier models already have context windows of around a million tokens, so on the input side Argon isn’t obviously ahead. In fact, Google hasn’t disclosed Argon’s context window at all. Output is where the gap is, and it matters because output is the scarcer resource for long-running work.

Why it matters for coding, research and agents

Output limits have quietly shaped how AI products are built. When a model can only write 128K tokens at a time, and some of that budget goes to its own thinking, developers have to break large jobs into chunks and stitch the results together. Each handoff risks losing context or introducing inconsistencies.

A 1M ceiling changes that calculation for several kinds of work.

In coding, large refactors, migrations and whole-module rewrites can be generated in one pass rather than file by file. Google says Argon agents are already migrating C and C++ code to Rust inside the company, at scales ranging from tens of thousands of lines to more than 800,000 lines for the Fuchsia Zircon kernel.

In research and knowledge work, long reports, legal drafting and multi-step financial analysis can run without being split up. Argon leads Google’s own comparison on enterprise benchmarks such as Vals Finance Agent v2 and Harvey’s Legal Agent Benchmark.

Agent workflows may benefit most. Agents that plan, act and reflect over long horizons use up output tokens quickly, so more headroom means fewer forced restarts.

The caveats

A bigger ceiling is not the same as better results. Google hasn’t said how well output quality holds up across hundreds of thousands of generated tokens, and long generations can accumulate errors that a shorter, checkpointed workflow might catch.

Cost is another factor. Argon launches at an introductory $2 per million input tokens and $10 per million output tokens, with cached input discounted by 95%. A full million-token response would cost about $10 at that rate. Google has said the price rises to $4 and $20 after the introductory period, which would put the same response at $20. It also hasn’t clarified whether reasoning tokens are billed at the output rate.

Availability is also limited for now. Argon is rolling out first to trusted cyber defenders through Google’s Fairwind Program, with paid API customers and Google AI Ultra subscribers next, and broader access after that. Most developers won’t be able to test the 1M limit yet.

Argon’s overall performance is competitive too. The model ties GPT-6 Astra on the Artificial Analysis Intelligence Index at a lower price, though it trails rivals on some terminal-driven agentic benchmarks. For OpenAI, the newer GPT-6.1 Sol lands just a point behind Astra on the same index, and Anthropic’s Claude Opus 5.5 and Claude Fable 5.1 remain the other models Argon is measured against.

The bottom line

The 1M output limit is a real, verified differentiator, and it targets a bottleneck that anyone building long-running agents or large-scale code generation tools will recognize. Whether it translates into noticeably better real-world results depends on how reliably Argon stays coherent over very long generations, and that will only become clear once the model reaches a wider set of developers. Expect OpenAI and Anthropic to face questions about whether 128K is still enough.

Posted in AI