Google launches three Gemini Flash AI models, delays flagship Pro
On July 21, 2026, Alphabet's Google unveiled a trio of new Gemini models aimed at cost‑efficient AI workloads. The lineup includes Gemini 3.6 Flash, the general‑purpose model; Gemini 3.5 Flash‑Lite, optimized for speed with a throughput of 350 output tokens per second; and Gemini 3.5 Flash‑Cyber, a cybersecurity‑focused variant that will initially be offered only to governments and trusted partners via CodeMender.
All three models reduce token consumption by about 17 % compared with the previous Gemini 3.5 Flash, with some benchmarks showing up to a 65 % cut, lowering per‑task costs. Pricing was announced at $1.50 per million input tokens and $7.50 per million output tokens for Flash‑6, with the Lite version priced even lower.
The anticipated flagship Gemini 3.5 Pro remains delayed; Google said it is still in testing and will be released when ready. In parallel, the company confirmed that pre‑training has begun on the next‑generation Gemini 4 model.