Google Releases Gemini 3.7 Flash, Competes With GPT 5.6 Terra & Muse Spark 1.2 On Benchmarks

Google DeepMind has rolled out Gemini 3.7 Flash, the latest update to its fast, lower-cost model line, positioning it as a direct competitor to OpenAI’s GPT-5.6 Terra and Muse’s Spark 1.2 across a wide range of benchmarks.

The release comes just weeks after Gemini 3.6 Flash, continuing Google’s pattern of rapid, incremental updates to the Flash series. Logan Kilpatrick, who leads product for Google’s AI Studio, announced the model on social media, highlighting its speed and calling out a “strong intelligence increase” delivered in roughly three weeks, which he attributed to algorithmic improvements from teams across Google DeepMind. He added that the company has been focused on making the model “feel more usable for real work,” and confirmed the model is now live in the API, AI Studio, Antigravity, and other surfaces.

Gemini 3.7 Flash Pricing

Alongside the performance bump, Google has cut pricing for the new model. Gemini 3.7 Flash is priced at $0.75 per million input tokens and $3.75 per million output tokens — the same introductory pricing as 3.6 Flash, which Google says represents a 50% reduction versus what 3.6 Flash will cost after the introductory window ends later this year. That undercuts both GPT-5.6 Terra ($2.00 input / $12.00 output) and Claude Sonnet 5 ($2.00 input / $10.00 output) by a wide margin, and even comes in cheaper than Muse Spark 1.2 ($1.25 input / $4.25 output).

Gemini 3.7 Flash Benchmark Performance

On Artificial Analysis’s Intelligence Index, a composite measure of model intelligence, Gemini 3.7 Flash scores 56 — ahead of Claude Sonnet 5 (55) and its own predecessor Gemini 3.6 Flash (52), though slightly behind GPT-5.6 Terra and Muse Spark 1.2, which both score 57.

gemini 3.7 flash benchmarks

The new model shows particularly strong results in coding and agentic tasks:

  • Code Arena (WebDev development, Elo): Gemini 3.7 Flash leads the pack at 1588, ahead of Muse Spark 1.2 (1535), Claude Sonnet 5 (1541), and GPT-5.6 Terra (1523).
  • FrontierCode 1.1 Main (production code quality): Gemini 3.7 Flash tops the field at 43.6%, versus 42.7% for Claude Sonnet 5 and 41.3% for GPT-5.6 Terra.
  • DeepSWE v1.1 (long-horizon software engineering): GPT-5.6 Terra leads here with 69.6%, followed by Muse Spark 1.2 (59.3%) and Gemini 3.7 Flash (65.3%), with Claude Sonnet 5 trailing at 54.0%.
  • Terminal-bench 2.1 and 3.0: GPT-5.6 Terra posts the top scores (87.4% and 20.8% respectively), with Gemini 3.7 Flash close behind on 2.1 (85.8%) but essentially tied with Claude Sonnet 5 on the newer 3.0 benchmark (14.9% vs 14.6%).

Gemini 3.7 Flash also posts big gains over its own predecessor in enterprise-oriented tests. On AutomationBench, which measures enterprise workflow automation, it scores 30.4% — nearly double Gemini 3.6 Flash’s 17.0%, and well ahead of GPT-5.6 Terra (23.6%) and Claude Sonnet 5 (10.7%). It also leads on Harvey LAB-AA (complex legal workflows) at 90.7% and GDP.PDF (expert PDF document comprehension) at 34.0%.

The model performs strongly on long-context and multimodal tasks as well, topping the field on LVBench (long video understanding, 85.4%) and GDM-MRCR v2 (long-context retrieval, 97.0% at 128k and 62.5% at 1M tokens).

Where Gemini 3.7 Flash falls behind is on some of the more demanding agentic and reasoning benchmarks. GPT-5.6 Terra leads on OSWorld-2.0 (agentic computer use, 50.2% vs Gemini’s 38.1%), while Claude Sonnet 5 tops Agent’s Last Exam (33.3% vs Gemini’s 26.3%) and BioMysteryBench‘s human-solvable category (87.5% vs 87.1%). On the harder “human difficult” tier of BioMysteryBench, GPT-5.6 Terra leads at 49.4%, with Gemini 3.7 Flash in second at 43.5%.

Muse Spark 1.2, meanwhile, posts the highest score on GDPVal-AA v2 (knowledge work, Elo 1628), ahead of Claude Sonnet 5 (1598) and GPT-5.6 Terra (1578), with Gemini 3.7 Flash trailing at 1525.

The Bigger Picture

The comparisons across GDPVal-AA v2 and AutomationBench suggest Google is positioning Flash increasingly as a model for enterprise and knowledge-work automation, not just a cheap, fast option for lightweight tasks. Kilpatrick’s post also noted the model’s rapid iteration cycle, going from 3.5 to 3.6 to 3.7 in quick succession, something he credited to “lots of hard work from the teams across GDM.”

With GPT-5.6 Terra and Muse Spark 1.2 still holding the edge on the top-line Artificial Analysis Intelligence Index and on the toughest agentic benchmarks, the race between the major labs’ fast, cost-efficient model tiers looks set to stay tight — with pricing now emerging as one of Google’s sharpest differentiators.

Posted in AI