Gemini 3.6 Flash Beats Gemini 3.5 Flash, Gemini 3.1 Pro On Most Benchmarks

Google’s new Flash model appears to be roughly at the same level as GPT 5.6 Luna and Grok 4.5.

Gemini 3.6 Flash, released alongside a smaller sibling called Gemini 3.5 Flash-Lite, clears its own predecessor and the older Gemini 3.1 Pro on every benchmark Google has published so far, while trading blows with rival mid-tier models from OpenAI and xAI depending on the task.

The timing is worth noting. Google had promised Gemini 3.5 Pro back at I/O for a June arrival, and that deadline has now slipped by weeks with no new date given. Instead of the flagship, Google has shipped two Flash-tier updates that slot in below where Pro is supposed to sit — and both come in cheaper than what they’re replacing. Gemini 3.6 Flash costs $1.50 per million input tokens and $7.50 per million output tokens, undercutting Gemini 3.5 Flash’s $9 output price despite scoring higher across the board. Gemini 3.5 Flash-Lite goes even further down, at $0.30 input and $2.50 output, positioned as the option for high-volume, cost-sensitive workloads.

Gemini 3.6 Flash Benchmarks

Against its direct predecessors, the result isn’t close. Gemini 3.6 Flash scores 58.7% on SWE-Bench Pro versus 55.1% for Gemini 3.5 Flash and 54.2% for Gemini 3.1 Pro. On DeepSWE v1.1, a long-horizon software engineering test, the gap widens considerably — Gemini 3.6 Flash hits 49% against 37% for 3.5 Flash and just 12% for 3.1 Pro. On Terminal-bench 2.1, the numbers land at 78.0%, 76.2%, and 73.8% respectively, and on MLE-Bench, Google’s new model scores 63.9% against 49.7% and 42.6% for the older two.

The pattern holds on knowledge work and multimodal tasks too. GDPVal-AA v2, which measures agentic performance on real-world economically valuable tasks and reports results as an Elo score, puts Gemini 3.6 Flash at 1421 against 1349 for 3.5 Flash and a much lower 965 for 3.1 Pro. On OSWorld-Verified, a computer-use benchmark, Gemini 3.6 Flash actually posts the best score in the entire comparison at 83.0%, ahead of every other model listed including GPT 5.6 Luna and Grok 4.5. Long context is where the gap turns lopsided: on GDM-MRCR v2 at 1M tokens (pointwise), Gemini 3.6 Flash scores 54.0% while Gemini 3.5 Flash and Gemini 3.1 Pro don’t clear 27%.

Where things get more interesting is the comparison against models outside Google’s own lineup. GPT 5.6 Luna leads on DeepSWE (67%) and Terminal-bench (84.7%), Grok 4.5 edges ahead on SWE-Bench Pro (64.7%), and Claude Sonnet 5 tops both MLE-Bench (66.9%) and GDPVal-AA v2 (1607). Gemini 3.6 Flash, meanwhile, holds the top spot on OSWorld-Verified, both CharXiv Reasoning variants, and both GDM-MRCR long-context tests. None of the four models sweeps the board, and the spread between them on most benchmarks is narrow enough that workload and price will likely decide which one gets picked for a given job more than raw scores will.

Google highlighted how Gemini 3.6 Flash was more token efficient than its predecessor.

gemini 3.6 flash token efficiency

Gemini 3.5 Flash-Lite Benchmarks

The smaller Flash-Lite model tells a similar story one tier down. Against Gemini 3.1 Flash-Lite, the March release it replaces, Gemini 3.5 Flash-Lite wins comfortably everywhere: 54.2% versus 38.3% on SWE-Bench Pro, 54.0% versus 31.0% on Terminal-bench 2.1, 39.2% versus 22.0% on MLE-Bench, and a GDPVal-AA v2 Elo of 1140 against 642. On OSWorld-Verified it scores 74.0% against 54.3%, and on the long-context GDM-MRCR tests it more than doubles its predecessor’s numbers at both context lengths.

Set against GPT-5.4 mini and Claude Haiku 4.5, the picture splits by task type. GPT-5.4 mini comes out ahead on SWE-Bench Pro (54.4%), Terminal-bench (59.2%), GDPVal-AA v2 (1171), and CharXiv Reasoning without tools (80.3%). Gemini 3.5 Flash-Lite answers back on MLE-Bench (39.2%, with GPT-5.4 mini not reporting a score), OSWorld-Verified (74.0% versus 72.1%), CharXiv with tools (76.5%), and both GDM-MRCR benchmarks by a wide margin. Claude Haiku 4.5 trails both models on nearly every metric in the table, despite carrying the highest price tag of the four at $1 input and $5 output.

Pricing is where Google’s pitch for the smaller model gets its edge. Gemini 3.5 Flash-Lite is a third of the input cost of GPT-5.4 mini and less than a third of Claude Haiku 4.5’s, while landing within a few points of both on most benchmarks and ahead on several. For workloads that run at high volume — content moderation, translation, bulk data processing — that combination of price and performance is likely to matter more than any single benchmark win.

Google still hasn’t given a new timeline for Gemini 3.5 Pro, and until it does, these two Flash-tier releases are effectively standing in for the flagship update the company promised back in May.

Posted in AI