Google shipped three new Gemini models on July 21, 2026 — Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber — but the release notably did not include the long-promised Gemini 3.5 Pro, which remains stuck in partner testing five months after its last update.
The Missing Model Overshadows the Launch
Rivals Have Not Been Waiting
The absence stands out because of how much ground competitors have covered in the meantime. Since Gemini Pro was last refreshed in February 2026, OpenAI has shipped GPT-5.5 and GPT-5.6, while Anthropic has released Claude Opus 4.8 and Claude Sonnet 5. Bloomberg has reported that Google has struggled to meet its own internal performance goals for the model, which explains part of the delay.
Logan Kilpatrick, who leads developer relations at Google, said the team is currently testing Gemini 3.5 Pro and hopes to "land soon," while confirming that pre-training has already begun on what he called the company's "most ambitious" Gemini 4 run yet. No timeline was given for either model.
Gemini 3.6 Flash Becomes the New Workhorse
Sharper Coding Scores, Lower Token Costs
Described by Google as its new workhorse model, Gemini 3.6 Flash cuts output token usage by up to 17% compared to 3.5 Flash, according to the Artificial Analysis Index. Coding benchmarks moved sharply: DeepSWE scores rose to 49% from 37%, MLE Bench climbed to 63.9% from 49.7%, and OSWorld-Verified reached 83.0% versus 78.4% for the previous version.
Google is pricing the model at $1.50 per million input tokens and $7.50 per million output tokens, and it is already live across Google AI Studio, Android Studio, Google Antigravity, the Gemini Enterprise Agent Platform, and the consumer Gemini app.
Flash-Lite Is Headed Straight Into Search
350 Tokens a Second at a Fraction of the Price
Gemini 3.5 Flash-Lite is pitched as the cheapest model in its class, generating 350 output tokens per second according to the Artificial Analysis Index. On Terminal-Bench 2.1 it scored 54%, up from 31% for the prior 3.1 Flash-Lite, and it posted 72.2% on GDM-MRCR v2 versus 60.1% before.
Pricing lands at $0.30 per million input tokens and $2.50 per million output tokens, and Google says the model is already rolling out inside Google Search — a deployment path that will expose it to a search user base far larger than any developer API ever could.
A Cybersecurity Model Kept Behind a Government-Only Wall
Flash Cyber Will Not Reach the Public
The third model, Gemini 3.5 Flash Cyber, is fine-tuned specifically to find and fix cybersecurity vulnerabilities and is said to perform competitively against the field's best on the CyberGym benchmark. Unlike the other two releases, it is not broadly available: Google is offering it exclusively to governments and trusted partners through its CodeMender platform, as part of a limited-access pilot program.
That restriction looks deliberate. A model capable of automatically scanning code for exploitable flaws and proposing fixes carries an obvious dual-use risk if it ends up in the wrong hands, and Google appears to be managing that risk by keeping distribution narrow for now.
Taken together, the announcement shows a company that can still move quickly at the low and mid tiers of its model lineup even as its flagship reasoning model slips further behind schedule. Whether Gemini 3.5 Pro arrives in time to answer GPT-5.6 and Claude Sonnet 5, or whether Google leapfrogs straight to Gemini 4, is now the more consequential question hanging over its AI roadmap.
