Gemini 3.8 Flash logo

Gemini 3.8 Flash

Google · Lançado 09/2026

Recomendado

Gemini 3.8 Flash is Google's strongest Flash workhorse yet — a genuine capability upgrade over the recommended 3.7 Flash at the same introductory price, with real gains across software engineering, agents and professional domains. The deciding caveat is Google's own: 3.8 "works harder," so per-task token burn runs well above 3.7, which remains the efficiency-first pick. Adopt 3.8 for long-horizon coding, agents and finance or legal work — benchmark your own workload before the intro price expires.

É adequada para ti?

Bom para

  • Long-horizon software engineering — DeepSWE v1.1 73.7% on Google's chart (74% ±1 on the independent leaderboard at high effort), up from 65.3% for 3.7 Flash
  • Autonomous agents and computer use — Terminal-Bench 4.0 19.1% vs 11.2% for 3.7, OSWorld 2.0 59.0% vs 50.6%
  • Knowledge-dense professional domains — HLE-Verified 54.9%; leads 3.7 and other frontier models on Vals Finance Agent V2 (61.4%) and Harvey's Legal Agent Benchmark
  • Google-ecosystem builders — stable GA ID gemini-3.8-flash in Gemini API, AI Studio, Antigravity, Android Studio, Gemini Enterprise; powers Gemini app, AI Mode and Gemini in Sheets

Não recomendado para

  • Token-metered, latency-sensitive production — "works harder" means extra reasoning steps and iterative tool calls; Artificial Analysis measured ~40% higher cost per task than 3.7 ($0.40 to $0.58)
  • Hardest frontier tasks — still trails Claude Opus 5 on Terminal-Bench 4.0 (19.1% vs 51.8%), OSWorld 2.0 (59.0% vs 75.4%) and GDP.PDF (35% vs 40% for GPT-5.6 Sol)
  • Non-English production without requalification — model card reports a 5.4-point multilingual safety regression vs 3.7 (lower is better); Google calls it mostly false positives
  • Teams needing price stability beyond 2026 — the introductory rate expires Dec 31, 2026, then doubles to $1.50/$7.50

Desempenho por tarefa

Software engineering (DeepSWE v1.1)

Very Good

73.7% Google-reported (74% ±1 independent at high effort) vs 65.3% for 3.7 Flash; outperforms most larger frontier models at a fraction of the cost

Terminal/CLI agent work (Terminal-Bench 2.1)

Very Good

89.4% vs 85.8% for 3.7 (Enterprise guide surface reports 90.8%)

General agentic work (Terminal-Bench 4.0)

Good

19.1% vs 11.2% for 3.7, but still far behind Claude Opus 5's 51.8%

Computer use (OSWorld 2.0)

Good

59.0% vs 50.6% for 3.7; behind Opus 5's 75.4%

Finance and legal agents (Vals Finance Agent V2 / Harvey Legal)

Very Good

61.4% and 10.0% all-pass respectively, both ahead of 3.7; Google says it leads other frontier models

Expert reasoning (HLE-Verified)

Very Good

54.9% vs 53.6% for 3.7

Prompt-injection robustness (Gray Swan)

Very Good

5.5% ASR@15, near Claude Opus 5's 4.8% and well ahead of 3.7 Flash's 9.2%

Preços

Entrada

$0.75 / 1M

Saída

$3.75 / 1M

Contexto

1M context, 64K output

Ver preços completos

Testes de desempenho

TestePontuaçãoFonte
DeepSWE v1.173.7% Fonte
DeepSWE v1.1 (independent leaderboard, high effort)74% ±1 Fonte
Terminal-Bench 2.189.4% Fonte
Terminal-Bench 4.019.1% Fonte
OSWorld 2.0 (batch tool)59.0% Fonte
HLE-Verified54.9% Fonte
Vals Finance Agent v261.4% Fonte
Gray Swan indirect prompt injection (ASR@15)5.5% Fonte

Ainda sem alterações de veredicto

O relógio começa a contar no primeiro dia: as alterações aparecem aqui à medida que o nosso veredicto evolui.

Registo de verificação

Ainda não há verificações

Ainda não registámos nenhuma verificação para esta entrada. Assim que uma for executada, o seu histórico aparece aqui.

Como avaliamos