Google Releases Gemini Flash 3.6 And Gemini Flash 3.5 Lite, Gemini Flash 3.6 Priced Cheaper Than 3.5 Flash

Google’s much-awaited Gemini 3.5 Pro is delayed, but’s gone ahead and launched two new models.

Gemini 3.6 Flash and Gemini 3.5 Flash Lite are now live in AI Studio and Vertex AI, giving developers a fresh pair of options to build with while the flagship Pro model remains stuck in internal testing.

Gemini 3.6 Flash

Gemini 3.6 Flash is priced at $1.50 per million input tokens and $7.50 per million output tokens, and carries a knowledge cutoff of March 2026 across all context lengths. Google is positioning it as the model that balances speed with intelligence for agentic and multimodal work, sitting in the same “flash” tier that Antigravity’s internal system had already slotted it into, above Flash Lite and below Pro. Interestingly, Gemini 3.6 Flash is priced cheaper than Gemini 3.5 Flash, which costs $9 per million output tokens compared to $7.5 for Gemini 3.6 Flash.

Google is saying that Gemini 3.6 Flash is much more token efficient than its predecessor. “This model is much more efficient and spend a lot less tokens to deliver better performance!” Google’s Logan Kilpatrick said on X.

Jeff Dean also shared a video showing off Gemini 3.6 Flash’s token efficiency.

Gemini 3.5 Flash Lite

Gemini 3.5 Flash Lite comes in considerably cheaper, at $0.30 per million input tokens and $2.50 per million output tokens, and is being billed as Google’s fastest, most cost-effective 3.5-series model for high-throughput execution. It shares the same March 2026 knowledge cutoff as its larger sibling. The pricing puts it roughly in line with where Gemini 3.1 Flash-Lite landed back in March, when it launched at $0.25 per million input tokens and $1.50 per million output tokens as Google’s cost-efficient option for high-volume workloads like translation and content moderation.

The Pro Problem

None of this changes the fact that Gemini 3.5 Pro, the model everyone’s actually been waiting on, is still nowhere to be found. Google had said on stage at I/O that Pro would arrive in June. That deadline came and went, and Bloomberg later reported the delay is tied to the model underperforming on coding benchmarks internally, with Google reportedly retraining on updated data in late June only to see disappointing results a second time. Google has since confirmed it’s testing 3.5 Pro with select partners and the US government, but there is still no public release date.

Leaning on a Flash release while Pro keeps cooking is a pattern Google has used before. When Gemini 3.5 Pro first missed its window, Google turned to Gemini 3.5 Flash, which went on to beat Gemini 3.1 Pro on several coding and agentic benchmarks despite being the cheaper, faster tier. Gemini 3.6 Flash appears built to run the same play. It doesn’t need to compete with GPT-5.6 or Claude Fable 5 at the very top of the leaderboard. It just needs to keep Google’s name in the conversation while Pro remains in limbo.

That conversation has gotten crowded fast. Google recently fell out of the top five labs on the Artificial Analysis Intelligence Index for the first time, with Gemini 3.5 Flash scoring 50 while Anthropic, OpenAI, SpaceXAI, Meta, and an open-source Chinese lab all ranked above it. In roughly a week, SpaceXAI’s Grok 4.5, three separate GPT-5.6 variants, Meta’s Muse Spark 1.1, and Moonshot AI’s Kimi K3 all launched, pushing the number of labs with a model scoring above 50 on the index from two to six.

What Developers Get

For teams building on the Gemini API, the practical shift is a new middle option between the ultra-cheap Flash Lite tier and whatever Pro eventually ships at. Gemini 3.6 Flash’s context window and thinking budget suggest Google is chasing the same agentic, long-horizon workloads that Gemini 3.5 Flash was built for back in May, where the model excelled at sub-agent deployment and multi-step coding cycles but came with a steep price jump over its predecessor. Whether 3.6 Flash repeats that cost jump or holds the line closer to where 3.1 Flash-Lite landed will matter a lot to teams running high-volume production traffic.

Google isn’t sitting still, but it is likely still sitting behind. Gemini 3.6 Flash won’t fix the Pro situation, and it won’t undo months of delay. What it does is put working, priced, generally-available models into developers’ hands today, while the model everyone’s actually waiting for stays somewhere inside Google’s testing pipeline.

Posted in AI