Gemini 3.6 Flash Scores 50 On Artificial Analysis Intelligence Index, Same As Gemini 3.5 Flash

Google has cut the price of Gemini 3.6 Flash by 17% compared to 3.5 Flash, but it doesn’t seem to have made much progress on the Artificial Analysis Intelligence Index.

Gemini 3.6 Flash lands at exactly 50 points on version 4.1 of the index, the identical score its predecessor was already sitting at, leaving Google’s best publicly available model in the same spot on the leaderboard it occupied before today’s launch. There is a takeaway though — Gemini 5.6 Flash, at $7.5 per million output tokens, is priced cheaper than Gemini 3.5 Flash, which cost $9.

Artificial Analysis rebuilt its Intelligence Index into version 4.1 last month, folding in nine evaluations including GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity’s Last Exam, GPQA Diamond, CritPt, AA-Omniscience, and AA-LCR. That rebuild is what dropped Google out of the top five AI labs for the first time, with Gemini 3.5 Flash scoring 50 while Anthropic, OpenAI, SpaceXAI, Meta, and an open-source Chinese lab all had models ranked above it. Gemini 3.6 Flash arriving at the same number confirms that the gap wasn’t a one-off measurement problem.

The top of the chart is currently led by Claude Fable 5 running with an Opus 4.8 fallback, at 60. It is followed by GPT-5.6 Sol at max effort leads at 59, followed by Kimi K3 at 57 and Claude Opus 4.8 at 56. GPT-5.6 Terra and GPT-5.5 sit tied at 55, Grok 4.5 follows at 54, and Claude Sonnet 5 comes in at 53. GPT-5.6 Luna, GLM-5.2, and Muse Spark 1.1 are bunched together at 51 apiece.

Gemini 3.6 Flash sits right below that cluster at 50, tied with its own predecessor and seven points clear of Gemini 3.1 Pro Preview, which trails at 46 alongside Qwen3.7 Max. That comparison matters more than the flat score against 3.5 Flash — Gemini 3.6 Flash is still beating Google’s own Pro-tier model from earlier this year, and it’s doing so at a lower price than the Flash model it replaces. Qwen3.7 Max, MiniMax-M3, DeepSeek V4 Pro, and MiMo-V2.5-Pro round out the bottom of the chart between 42 and 46, all open-weights models that continue to trail the closed labs on this particular index.

Artificial Analysis’s own testing, published ahead of release, confirms the intelligence score is going nowhere while everything around it moves. Gemini 3.6 Flash matches Gemini 3.5 Flash’s Intelligence Index score almost evaluation for evaluation, with the two exceptions being a 72-point jump on GDPval-AA v2, to an Elo of 1421, and a three-point slide on Humanity’s Last Exam, down to 38%. On AA-Briefcase, Artificial Analysis’s agentic knowledge-work benchmark, Gemini 3.6 Flash climbs 95 points to an Elo of 961. None of that moves the composite score, but it does suggest the model is getting better at doing agentic work even while its raw reasoning number holds flat.

Where Gemini 3.6 Flash does pull ahead is speed and cost. Average time per task drops to 1.3 minutes from 2.7 for Gemini 3.5 Flash, a reduction of more than half, with output measured at 304 tokens per second in Artificial Analysis’s pre-launch runs. Cost per task falls from $0.59 to $0.50, roughly 18% cheaper, tracking the new $1.50/$7.50 pricing against the old $1.50/$9.00. The model is doing the same job in less time, at lower cost, without a matching jump in the headline number that gets it noticed.

Gemini 3.5 Flash-Lite tells the opposite story. It scores 36 on the Intelligence Index, an 11-point jump over Gemini 3.1 Flash-Lite’s 25, putting it behind Nemotron 3 Ultra (38) and DeepSeek V4 Flash at max effort (40), and ahead of Mistral Medium 3.5 (30). The gains are concentrated in agentic tasks — GDPval-AA v2 jumps 498 points to an Elo of 1140, AA-Briefcase climbs 413 points to 634, Terminal-Bench v2.1 improves by 22.5 points to 53.6%, and τ³-Banking gains 7.8 points to 16.5%. Time per task nearly halves too, down to 0.6 minutes from 1.0, with output speed measured at 350 tokens per second. The catch is price: cost per task more than doubles, from $0.04 to $0.09, on new pricing of $0.30 input and $2.50 output against the previous $0.25 and $1.50 — an increase that lands despite the model actually using fewer output tokens per task, averaging 13,000 against 20,000 for its predecessor.

Both models keep the same 1 million token context window and multimodal input support — text, image, video, and speech in, text out — as the versions they replace, and both retain the standard 90% discount on cached input tokens.

What the flat score on Gemini 3.6 Flash really points to is a company running out of easy wins at the Flash tier while its actual flagship stays stuck in delay. Gemini 3.5 Pro was supposed to ship in June, and Google has offered no new date since it missed that window. Until it does, Gemini 3.6 Flash’s job isn’t to climb the leaderboard — it’s to get the same work done faster and cheaper while the intelligence number sits still, and let Gemini 3.5 Flash-Lite pick up the slack lower down the stack.

Posted in AI