AI Models & Platforms
Google Ships Three Gemini Flash Models as Its Flagship Slips

Google released three new Gemini models on July 21, 2026 — Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and a security-tuned model called Gemini 3.5 Flash Cyber — while the flagship it once promised for June, Gemini 3.5 Pro, stays in limited testing with partners. In the same update, the company said it has already begun what it calls “our most ambitious pre-training run yet” for Gemini 4.
All three new models sit at the cheaper, faster end of Google’s lineup: the “Flash” tier built for speed and high volume rather than maximum reasoning depth. That framing matters because the model meant to compete at the top of the field, 3.5 Pro, has slipped well past the target Google set on stage at its I/O conference in May 2026, where it said the Pro version would arrive the following month. Google now says only that Pro is testing with partners and will ship when it is ready.
What the new models do
Gemini 3.6 Flash is the direct successor to the 3.5 Flash model Google launched at I/O. Its main selling point is efficiency: Google says it uses about 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index and takes fewer reasoning steps and tool calls to finish multi-step jobs. It is priced at $1.50 per million input tokens and $7.50 per million output tokens, cheaper on output than the $9 the previous Flash charged. Its knowledge cutoff moves forward more than a year, from January 2025 to March 2026.
On Google’s own coding and agent benchmarks, 3.6 Flash outscores its predecessor — 49% versus 37% on the DeepSWE coding test, and 83% versus 78.4% on the OSWorld computer-use benchmark, among others. Those are vendor-reported numbers, not independent measurements, and no outside evaluator has replicated them yet. They describe how the new model compares to the old one, not where it lands against rival flagships.
Gemini 3.5 Flash-Lite is the cheaper companion, aimed at high-throughput, low-latency work such as agentic search and document processing. Google prices it at $0.30 per million input tokens and $2.50 per million output tokens and says it is a clear step up from the 3.1 Flash-Lite model it replaces. Both new models are live in the Gemini app today, with Flash-Lite also coming to Search, and developers can reach them through Google’s Antigravity, AI Studio, and Android Studio tools.
A security model kept on a short leash
The third release is the more unusual one. Gemini 3.5 Flash Cyber is fine-tuned to find, validate, and patch software vulnerabilities, and it runs inside CodeMender, Google DeepMind’s code-security agent. Rather than open it up, Google is restricting the model to governments and trusted partners under what it describes as a limited-access pilot.
CodeMender is the reason a cheap model makes sense here. The agent scans a codebase for flaws, works out the root cause of each one, then generates and self-checks a patch before routing only high-quality fixes to a developer for approval — an approach Google DeepMind said had already produced 72 upstreamed fixes to open-source projects over its first six months. Running that loop with a Flash-tier model rather than a large flagship is what lets Google pitch vulnerability detection at scale and at a lower price per token than bigger models.
The gating is the tell that this is dual-use technology. A model good at finding exploitable bugs is equally useful to an attacker, and Google framed the restricted release around that tension, saying the model “will give frontline defenders a head start in finding and fixing critical vulnerabilities before they can be exploited, while mitigating against broader misuse.” It is the same calculation showing up across autonomous AI security agents: the capability that helps defenders is the capability that worries them.
The flagship gap
None of this closes the gap that has defined Google’s summer. The Flash releases are incremental, and the model that would actually contest the top of the market is still absent while rivals ship. In roughly a week, Grok 4.5, three variants of OpenAI’s GPT-5.6, and Moonshot AI’s Kimi K3 all launched, and Anthropic’s Claude family, led by its Fable 5 model, sits atop the public leaderboards Google’s Flash tier now trails. The pressure is not only competitive: Google is also fending off regulators in Europe, who ordered it to share Android and Search data with rival AI assistants.
Google’s answer, for now, is to compete on price and speed at the Flash tier rather than on top-line capability — the same play it ran in May, when 3.5 Flash beat the older 3.1 Pro on several coding tests despite being the cheaper option. It keeps Google in the conversation without answering the question developers are actually asking.
That is the backdrop against which the Gemini 4 line should be read. Starting a pre-training run is a statement of intent, not a shipped capability, and it says nothing about what the resulting model will do until the evaluations exist. It is notable mainly for what it signals: Google is talking up the next generation while the current one’s flagship remains stuck in testing. The measures that matter now are whether 3.5 Pro finally ships, and whether Gemini 4 turns out to be more than a training run.












