Gemini 3.8 Flash logo

Gemini 3.8 Flash

Google · Lanzado sept 2026

Recomendado

Gemini 3.8 Flash is Google's strongest Flash workhorse yet — a genuine capability upgrade over the recommended 3.7 Flash at the same introductory price, with real gains across software engineering, agents and professional domains. The deciding caveat is Google's own: 3.8 "works harder," so per-task token burn runs well above 3.7, which remains the efficiency-first pick. Adopt 3.8 for long-horizon coding, agents and finance or legal work — benchmark your own workload before the intro price expires.

¿Es adecuada para ti?

Bueno para

  • Long-horizon software engineering — DeepSWE v1.1 73.7% on Google's chart (74% ±1 on the independent leaderboard at high effort), up from 65.3% for 3.7 Flash
  • Autonomous agents and computer use — Terminal-Bench 4.0 19.1% vs 11.2% for 3.7, OSWorld 2.0 59.0% vs 50.6%
  • Knowledge-dense professional domains — HLE-Verified 54.9%; leads 3.7 and other frontier models on Vals Finance Agent V2 (61.4%) and Harvey's Legal Agent Benchmark
  • Google-ecosystem builders — stable GA ID gemini-3.8-flash in Gemini API, AI Studio, Antigravity, Android Studio, Gemini Enterprise; powers Gemini app, AI Mode and Gemini in Sheets

No recomendado para

  • Token-metered, latency-sensitive production — "works harder" means extra reasoning steps and iterative tool calls; Artificial Analysis measured ~40% higher cost per task than 3.7 ($0.40 to $0.58)
  • Hardest frontier tasks — still trails Claude Opus 5 on Terminal-Bench 4.0 (19.1% vs 51.8%), OSWorld 2.0 (59.0% vs 75.4%) and GDP.PDF (35% vs 40% for GPT-5.6 Sol)
  • Non-English production without requalification — model card reports a 5.4-point multilingual safety regression vs 3.7 (lower is better); Google calls it mostly false positives
  • Teams needing price stability beyond 2026 — the introductory rate expires Dec 31, 2026, then doubles to $1.50/$7.50

Rendimiento por tarea

Software engineering (DeepSWE v1.1)

Very Good

73.7% Google-reported (74% ±1 independent at high effort) vs 65.3% for 3.7 Flash; outperforms most larger frontier models at a fraction of the cost

Terminal/CLI agent work (Terminal-Bench 2.1)

Very Good

89.4% vs 85.8% for 3.7 (Enterprise guide surface reports 90.8%)

General agentic work (Terminal-Bench 4.0)

Good

19.1% vs 11.2% for 3.7, but still far behind Claude Opus 5's 51.8%

Computer use (OSWorld 2.0)

Good

59.0% vs 50.6% for 3.7; behind Opus 5's 75.4%

Finance and legal agents (Vals Finance Agent V2 / Harvey Legal)

Very Good

61.4% and 10.0% all-pass respectively, both ahead of 3.7; Google says it leads other frontier models

Expert reasoning (HLE-Verified)

Very Good

54.9% vs 53.6% for 3.7

Prompt-injection robustness (Gray Swan)

Very Good

5.5% ASR@15, near Claude Opus 5's 4.8% and well ahead of 3.7 Flash's 9.2%

Precios

Entrada

$0.75 / 1M

Salida

$3.75 / 1M

Contexto

1M context, 64K output

Ver precios completos

Pruebas de rendimiento

PruebaPuntuaciónFuente
DeepSWE v1.173.7% Fuente
DeepSWE v1.1 (independent leaderboard, high effort)74% ±1 Fuente
Terminal-Bench 2.189.4% Fuente
Terminal-Bench 4.019.1% Fuente
OSWorld 2.0 (batch tool)59.0% Fuente
HLE-Verified54.9% Fuente
Vals Finance Agent v261.4% Fuente
Gray Swan indirect prompt injection (ASR@15)5.5% Fuente

Aún no hay cambios de veredicto

El reloj corre desde el primer día: los cambios aparecerán aquí a medida que evolucione nuestro veredicto.

Registro de verificación

Aún no hay comprobaciones de verificación

Todavía no hemos registrado ninguna comprobación de verificación para esta entrada. Cuando se ejecute una, su historial aparecerá aquí.

Cómo evaluamos