Google is back in the top 3 in the AI model race.
Google’s new Gemini 4 Argon has matched OpenAI’s flagship on one of the industry’s most closely watched independent scorecards, and it’s doing so at a lower cost per task. Artificial Analysis says Argon scores 53 on its Intelligence Index, equal to GPT-6 Astra (max), while costing about 60% as much to run at the model’s current discounted pricing. The firm says the result puts Google back among the top three labs in intelligence.
Argon is the first Google proprietary model above the Flash class in more than seven months, and Artificial Analysis tested it at high reasoning, the highest setting currently available. The model is being rolled out to selected users and isn’t publicly available yet.

Where Argon Lands On The Intelligence Index
On the latest version of the index (v4.3.2, which combines 10 evaluations including AA-Briefcase, GDPval-AA, AutomationBench-AA, Terminal-Bench 4.0, Humanity’s Last Exam and CritPt), Argon at 53 sits level with GPT-6 Astra (max) and Claude Fable 5.1 (max with fallback). It is one point ahead of GPT-6.1 Sol (max) at 52.
The jump from Google’s earlier models is large. Argon is 23 points above Gemini 3.1 Pro Preview, which scored 30, and 12 points above Gemini 3.8 Flash (high) at 41. Artificial Analysis attributes the gains to lower hallucination rates and stronger agentic capabilities.
Anthropic still holds the top of the chart: Claude Opus 5.5 (max with fallback) scores 58 and Claude Sonnet 5.5 (max with fallback) scores 56. Argon ties for third place in a group at 53, ahead of Meta’s Muse Spark 1.3 (max) at 48 and xAI’s Grok 4.7 (xhigh) at 46. That makes Argon a clear step up for Google.
Gemini 4 Argon Pricing
Pricing is where Argon’s headline advantage comes from. At standard pricing of $4 per million input tokens and $20 per million output tokens, Google is currently running a 50% promotion that brings the rate down to $2 and $10. Cached input tokens get a 95% discount, which works out to $0.10 per million tokens at the discounted price, up from a 90% cached discount on Gemini 3.8 Flash.
At the discounted rate, Argon costs $1.99 per Intelligence Index task. GPT-6 Astra (max) costs $3.26 for the same measure, so Argon comes in at roughly 60% of Astra’s cost, or about 40% cheaper. Cheaper tokens are doing all the work here: Argon averages 62,000 output tokens per task against 27,000 for Astra, so it is actually far more verbose.

There are two important caveats. First, Google hasn’t confirmed when the promotion ends. Once standard pricing applies, Argon’s cost per task rises to $3.98, around 1.2 times Astra’s. Second, OpenAI’s own cheaper tier still beats Argon on cost: GPT-6.1 Sol (max) scores one point lower at 52, but Argon at the discounted rate costs about 2.7 times as much per task. Argon doesn’t sit in the chart’s most attractive quadrant for intelligence per dollar, even with the discount in place.
Stronger Agentic Performance
Agentic work has historically been a weak spot for Gemini models, and Artificial Analysis says that’s changing. Argon takes the top spot on AutomationBench-AA, the Zapier-developed benchmark that Artificial Analysis runs independently, with 77.5%. That’s six points ahead of Claude Sonnet 5.5 (max) at 71.3% and well ahead of Claude Opus 5.5 at 69.5%, GPT-6 Astra at 68.5% and Claude Fable 5.1 at 59.4%.

On Terminal-Bench 4.0, Argon scores 57.1%, an improvement of 53 points over Gemini 3.1 Pro Preview’s 4.0%. That still leaves it behind Claude Sonnet 5.5 (63.6%), Claude Opus 5.5 (59.6%) and GPT-6 Astra (59.1%), though it is ahead of GPT-6.1 Sol (56.1%) and Fable 5.1 (52.0%). Terminal-based tasks were also where Argon trailed rivals in Google’s own benchmark table at launch.
On AA-Briefcase, Artificial Analysis’ agentic knowledge work benchmark, Argon scores 1494 Elo. The result is driven by a 65% rubric pass rate, the highest the firm has recorded, but is held back by lower Analytical Quality (1576 Elo) and Presentation Quality (1308 Elo). In other words, Argon tends to hit the checklist, but its outputs are less polished than the best models.
Lowest Hallucination Rate Among Leading Models
The most striking number in the analysis may be on AA-Omniscience, which measures knowledge and hallucination. Argon has a 15% hallucination rate, the lowest of any model scoring 45 or above on the Intelligence Index. By comparison, GPT-6 Astra (max) hallucinates 51% of the time and GPT-6.1 Sol (max) 54%. In practice, that means Argon is much more likely to say it doesn’t know something rather than guess wrong.
The tradeoff is accuracy. Argon scores 50%, five points lower than Gemini 3.1 Pro Preview and 13 points below Astra’s 63%. Its overall AA-Omniscience score of 42 is essentially in line with Astra (43) and Sol (42).

Model Details
Argon has a 1 million token context window and accepts text, image, video and speech input, with text output. Artificial Analysis also tested the model with Long Decode Continuation, a new Gemini API feature that pauses long responses and resumes them across follow-up calls, allowing reasoning to run up to 1 million output tokens without request timeouts.
What It Means
Argon’s Intelligence Index result confirms that Google has closed much of the gap at the top end of the market after a long stretch where its best models were Flash-class. Matching Astra and Fable 5.1 at 53 is a meaningful achievement, and the low hallucination rate and AutomationBench lead give it concrete strengths. But the cost advantage is tied to a temporary discount, Anthropic’s newest models still lead the index, and Argon isn’t yet available for independent users to test for themselves. How Google prices the model once the promotion ends will determine whether the efficiency story holds.