Jul 22, 2026 · 1:25 AM
Subscribe
Home Ai

Google released three Gemini models in one day while its flagship is still stuck in testing

Google released Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber on July 21, 2026, cutting output pricing to $7.50 per million tokens and entering enterprise security tooling with a vulnerability-finding model. Meanwhile, Gemini 3.5 Pro, originally due in June, has missed three straight deadlines over coding performance issues.

Walter Schulze
· 5 min read · 542 reads
Google released three Gemini models in one day while its flagship is still stuck in testing

Google shipped three Gemini models on July 21, but the model everyone is waiting for still isn't here. The Flash releases give developers cheaper tools today, while Gemini 3.5 Pro remains the unanswered question.

Three new models in one day looks like confidence. Read it closely and it also looks like pressure. Google announced Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber on July 21, cutting output prices, talking up faster agent workflows, and moving deeper into security while Gemini 3.5 Pro, the flagship model promised for June after Google I/O, is still being tested with partners.

That is the story. Google can ship. It just hasn't shipped the one model meant to prove it can still lead at the top of the market.

Gemini 3.6 Flash is the main release for most developers. According to Google's own announcement, the model reduces output token use by 17% compared with Gemini 3.5 Flash on the Artificial Analysis Index, and Google says it takes fewer reasoning steps and tool calls on multi-step workflows. The price moved as well: input stays at $1.50 per million tokens, while output drops from $9 to $7.50 per million tokens.

Those numbers are not decorative. If you're building agents that call a model again and again, token use is the bill. A 17% reduction in output tokens does more work for a team than another vague claim about intelligence. Artificial Analysis also found that Gemini 3.6 Flash cut average time per task from 2.7 minutes to 1.3 minutes compared with 3.5 Flash, while its Intelligence Index score stayed flat at 50. Faster, not much smarter. That is still useful.

Google is positioning 3.6 Flash as its workhorse model for coding, knowledge work, and multimodal tasks - the everyday developer workload, essentially. The company pointed to DeepSWE gains of 49% versus 37% for 3.5 Flash, plus MLE Bench gains of 63.9% versus 49.7%. You don't need to pretend every benchmark settles the argument. The practical point is simpler: Google is trying to make the Flash tier less chatty, cheaper to run, and better at the kinds of repeated tool calls developers actually pay for.

Google is chasing the high-volume market

Flash-Lite is the cheaper, faster sibling. Google says Gemini 3.5 Flash-Lite runs at 350 output tokens per second, as measured by Artificial Analysis, and prices it at $0.30 per million input tokens and $2.50 per million output tokens. It is available through the Gemini API via Google AI Studio and Android Studio, and Google says it is also rolling out in Search.

This is where Google's move is strongest. These aren't the models that win the loudest benchmark headlines. They're the ones that let a company run document processing, agentic search, receipt translation, or customer workflow automation without watching inference costs eat the product margin. If you're a startup trying to build on top of Gemini, that matters more than a leaderboard screenshot.

Gemini 3.5 Flash Cyber is the sharper release. Google says the model is built on 3.5 Flash and fine-tuned for finding, validating, and patching vulnerabilities inside CodeMender. The Verge reported that, in Google's V8 JavaScript engine tests, the Cyber model found 55 confirmed issues, compared with 47 for standard Gemini 3.5 Flash and 36 for Anthropic's Opus 4.6. It also found 10 novel issues that other models missed.

Take vendor benchmarks with care. Still, the signal is plain enough. Google is not just adding another general-purpose model to a price sheet. It is trying to put Gemini into enterprise security work, where a model's value is measured in bugs found, patches proposed, and time saved by teams that already have too much code to review.

The rollout is restricted. Google says 3.5 Flash Cyber will be available soon to governments and trusted partners through a limited pilot because the same technology that helps defenders can also help attackers. That caution is not a footnote. It is the product boundary.

The Pro delay is still the problem

Frankly, the more revealing news is what didn't ship. Gemini 3.5 Pro was announced around Google I/O in May, when Google said it expected to roll out the model the following month. June passed. As 9to5Google noted from Bloomberg's reporting, Google has been trying to improve the model's capabilities, especially coding, and a late-June training data update reportedly produced disappointing results.

Google's public line is narrower than the rumors around it. In its July 21 post, the company said Gemini 3.5 Pro is testing with partners and will be made broadly available when it is ready. It also said it has started its most ambitious pre-training run yet for Gemini 4. That is a useful signal. It is not a launch date.

Here is the thing: the Flash tier can be very good and still leave Google exposed. Flash handles cost, latency, and volume. Pro is supposed to carry the frontier story. OpenAI, Anthropic, xAI, Meta, Moonshot, you name it, are all trying to define that top end with models that can code, reason, and run long workflows with fewer obvious failures.

Google's July 21 release gives developers real tools now. That should not be dismissed. 3.6 Flash is cheaper on output, Flash-Lite is built for raw throughput - and Flash Cyber gives Google a credible security product with a controlled rollout.

But three mid-tier wins don't erase one missing flagship. Until Gemini 3.5 Pro arrives, Google's AI story has a split screen: useful shipping on one side, unfinished proof on the other.

Also read: Intel cuts jobs in its fastest-growing division two days before Q2 earningsSupermicro books $60 billion in orders in a single quarter as AI server backlog hits record highOpenAI adds two bank CEOs to its board as IPO signals grow stronger

TOPICS
Walter Schulze brings all the breaking news stories in the tech and startup world and to ensure that Startup Fortune offers a timely reporting on the trends happen in the industry. He now works on a part time basis for Startup Fortune specializing in covering tech and startup news and he also sheds light on investment opportunities and trends.
Related Articles
More posts →
Loading next article…
You're all caught up