Three Gemini Flash Models Ship — 3.5 Pro Still Delayed

Google DeepMind logoGoogle DeepMindImportante22 de julio de 2026Modelos
Qué pasó
Google shipped Gemini 3.6 Flash ($1.50/$7.50 per 1M tokens with 49% DeepSWE), 3.5 Flash-Lite ($0.30/$2.50 at 350 tok/s), and Flash Cyber (government-gated cybersecurity model) — three efficiency-focused models optimized for agent workloads.
Por qué importa
Google is executing on cost and speed but cannot ship a frontier model — Gemini 3.5 Pro has missed its June release target, is reportedly months behind on coding benchmarks, and the competitive gap at the top is widening while Gemini 4 pretraining begins.
Qué hacer
Benchmark 3.6 Flash on your agent workloads — the 17% token reduction compounds at scale — and evaluate Flash-Lite for high-throughput pipelines, but build on what exists while 3.5 Pro remains unshipped.

Google DeepMind shipped three Gemini Flash models on Tuesday — and not one of them is the model everyone's been waiting for.

Gemini 3.6 Flash ($1.50/$7.50 per 1M tokens) is the new workhorse: 17% fewer output tokens than 3.5 Flash, a 49% DeepSWE score, and 63.9% on MLE-Bench. Gemini 3.5 Flash-Lite ($0.30/$2.50) hits 350 tokens per second — 2x faster than its predecessor — and nearly matches Gemini 3 Flash on coding benchmarks at a fraction of the cost. Gemini 3.5 Flash Cyber is a cybersecurity-specialized model, available exclusively to governments and trusted partners via CodeMender (Google Blog, 2026).

But the flagship Gemini 3.5 Pro — announced at I/O in May — has now missed multiple release targets. Google product lead Logan Kilpatrick says it's "testing with partners." Bloomberg reports the model is months behind schedule on coding benchmarks after struggling to meet internal performance goals. Meanwhile, Gemini 4 pretraining has begun (Google Blog, 2026; TechCrunch, 2026).

What happened

Gemini 3.6 Flash is the direct successor to 3.5 Flash. At $1.50/M input and $7.50/M output tokens with a 1M-token context window, it delivers the same capabilities while using 17% fewer output tokens — a direct attack on the token-bloat problem inflating agent costs.

BenchmarkGemini 3.5 FlashGemini 3.6 FlashImprovement
DeepSWE37%49%+12 points
MLE-Bench49.7%63.9%+14.2 points
OSWorld-Verified78.4%83.0%+4.6 points

Gemini 3.5 Flash-Lite hits 350 tokens per second — 2x faster than 3.1 Flash-Lite — while scoring 54.2% on SWE-Bench Pro and 74.0% on OSWorld. At $0.30/M input and $2.50/M output, it's priced between DeepSeek V4 Flash and Pro on input cost (VentureBeat, 2026).

Gemini 3.5 Flash Cyber is fine-tuned for vulnerability hunting and patching. Available exclusively via CodeMender to governments and trusted partners. Competitive with Anthropic's Mythos on CyberGym.

Gemini 3.5 Pro (announced I/O May 2026): repeatedly delayed past its promised June launch. "Testing with partners, hope to land soon" per Google's Logan Kilpatrick. Gemini 4 pretraining has begun.

Why it matters

Google is winning on cost and speed but can't ship a frontier model. The efficiency gap between Google's Flash tier and its competitors is widening — 3.6 Flash at $1.50/$7.50 with 49% DeepSWE is the best price-to-performance ratio Google has ever shipped. But the Pro gap at the top is structural: repeated delays, missed internal performance goals (per Bloomberg), and no confirmed timeline.

The AI race shifted from benchmark scores to cost efficiency. The models that win agent workloads aren't the ones with the highest benchmark scores — they're the ones with the best performance per dollar. Google understands this. The Flash line is optimized for exactly that metric.

What changes for you

  • Benchmark Gemini 3.6 Flash on your agent workflows — 17% fewer output tokens and 49% DeepSWE at $1.50/M input is a meaningful improvement
  • Evaluate Flash-Lite for high-throughput pipelines — 350 tok/s at $0.30/M input, 2x faster than the prior generation
  • Do not plan architecture around Gemini 3.5 Pro — it is repeatedly delayed with no confirmed timeline. Build for what exists
  • Flash Cyber is available only to governments and trusted partners — not a commercial option

FAQ

Is Gemini 3.6 Flash a drop-in replacement for 3.5 Flash? Yes. Same API, same 1M context window, but 17% fewer output tokens on the same tasks.

When will Gemini 3.5 Pro ship? Unknown. Google says "soon" but has missed multiple prior deadlines. The model is behind on coding benchmarks.

Can I try Flash Cyber? No — it's exclusively available to governments and trusted partners via CodeMender.

Herramientas y modelos afectados

No vuelvas a tener que ponerte al día

El resumen semanal — solo cambios de veredicto y acciones urgentes. Sin relleno.

Al suscribirte aceptas nuestra Política de privacidad. Cancela cuando quieras.