Google Ships Three New Gemini Models, but the Flagship Pro Is Still Missing

Google DeepMind released three new Gemini models on July 21, 2026: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber. The top-tier Gemini 3.5 Pro, which Google had said was coming two months earlier, was not part of the release.
Google described Gemini 3.6 Flash as its "workhorse model," pointing to improvements in coding, knowledge work, and multimodal performance (meaning it handles text, images, and other input types together) over the previous 3.5 Flash. It also uses up to 17% fewer tokens than its predecessor, according to Google's announcement on its official blog (TechCrunch; Google Blog). Tokens are the chunks of text an AI model processes; fewer tokens per task means lower costs for developers. Google said 3.6 Flash builds directly on developer and customer feedback gathered since the 3.5 Flash release.
Gemini 3.5 Flash-Lite is positioned as the most cost-effective model in its class. Gemini 3.5 Flash Cyber, the third model in the release, is fine-tuned specifically for finding and fixing cybersecurity vulnerabilities. Its availability is restricted: Google stated it will be offered exclusively to governments and trusted partners through a limited-access pilot program (TechCrunch).
Google framed the July 2026 releases as focused on delivering efficiency, latency (how quickly a model responds), and reliability to customers building AI agents at scale. That framing aligns with a broader industry shift: buyers are moving away from raw benchmark scores toward what it actually costs to run a model in production and how dependably it performs under real workloads.
The conspicuous absence from the release is Gemini 3.5 Pro. Google last updated the Pro tier in February 2026. During the Gemini 3.5 Flash launch in May 2026, the company said 3.5 Pro was already in internal use and would roll out the following month. That did not happen. Bloomberg reported on July 16, 2026 that Google was facing internal delays in launching 3.5 Pro as it struggled to meet its own performance targets (TechCrunch).
Logan Kilpatrick, Google DeepMind's product lead, stated on July 21 that the company is currently testing Gemini 3.5 Pro with partners and hopes to "land soon." He also disclosed that the team has begun what he called its most ambitious pre-training run yet for Gemini 4 (TechCrunch). Pre-training is the initial, resource-intensive phase where a model learns from massive datasets before being refined for specific tasks.
The Pro delay carries practical weight for paying customers. Google's AI Ultra plan, priced at $99.99 per month as of July 2026, provides access to Gemini 3.1 Pro with advanced research features (Reuters). Subscribers on that tier have now gone five months without a Pro-tier update, even as the Flash line has received two successive generations in that window.
A separate Reuters report dated July 20, 2026, citing The Information, indicated that Google is developing a new server chip designed to run Gemini models more efficiently (Reuters). That effort sits alongside the model-level efficiency gains claimed for 3.6 Flash, suggesting Google is pressing on inference cost reduction across both the silicon and model layers simultaneously.
The Flash Cyber model warrants attention on its own. Domain-specific models fine-tuned for security work are not entirely new, but restricting access to governments and trusted partners marks a deliberate distribution choice. Google is clearly keeping the model out of general developer access, which limits scrutiny but also limits misuse potential. Whether that tradeoff holds up depends on how the pilot program defines "trusted partners" and what transparency, if any, accompanies the model's deployment. Google did not provide those details in its announcement.
The broader context here is that Google is now running a two-track cadence: rapid iteration on Flash-tier models for cost-sensitive, high-volume workloads, and a slower, more deliberate path on the Pro tier. The 17% token reduction in 3.6 Flash is the kind of metric that matters directly to developers pricing agent pipelines. Whether that efficiency gain holds across real-world workloads, as opposed to Google's internal benchmarks, will be tested quickly by the developer community.
The mention of Gemini 4 pre-training, even in passing, signals that Google is not slowing its roadmap pace despite the Pro-tier bottleneck. In my view, if 3.5 Pro's delays reflect genuine performance gaps that Google is unwilling to ship around, that is arguably the right call. Shipping a flagship model that underperforms its own internal expectations would damage more than a delay does. But the gap between the May promise and the July reality is now wide enough that Kilpatrick's "land soon" carries more weight than a typical product-leader hedge.
For developers and enterprises building on Gemini today, the practical takeaway is straightforward. The Flash tier continues to improve at a brisk pace, with tangible cost and efficiency gains. The Pro tier remains frozen at February's 3.1 release, with no firm ship date. And somewhere in Google's infrastructure, Gemini 4 is already in pre-training.


