Gemini 3.8 Flash logo

Gemini 3.8 Flash

Google · Released Sep 2026

Recommended

Gemini 3.8 Flash is Google's strongest Flash workhorse yet — a genuine capability upgrade over the recommended 3.7 Flash at the same introductory price, with real gains across software engineering, agents and professional domains. The deciding caveat is Google's own: 3.8 "works harder," so per-task token burn runs well above 3.7, which remains the efficiency-first pick. Adopt 3.8 for long-horizon coding, agents and finance or legal work — benchmark your own workload before the intro price expires.

Is it right for you?

Good for

  • Long-horizon software engineering — DeepSWE v1.1 73.7% on Google's chart (74% ±1 on the independent leaderboard at high effort), up from 65.3% for 3.7 Flash
  • Autonomous agents and computer use — Terminal-Bench 4.0 19.1% vs 11.2% for 3.7, OSWorld 2.0 59.0% vs 50.6%
  • Knowledge-dense professional domains — HLE-Verified 54.9%; leads 3.7 and other frontier models on Vals Finance Agent V2 (61.4%) and Harvey's Legal Agent Benchmark
  • Google-ecosystem builders — stable GA ID gemini-3.8-flash in Gemini API, AI Studio, Antigravity, Android Studio, Gemini Enterprise; powers Gemini app, AI Mode and Gemini in Sheets

Not good for

  • Token-metered, latency-sensitive production — "works harder" means extra reasoning steps and iterative tool calls; Artificial Analysis measured ~40% higher cost per task than 3.7 ($0.40 to $0.58)
  • Hardest frontier tasks — still trails Claude Opus 5 on Terminal-Bench 4.0 (19.1% vs 51.8%), OSWorld 2.0 (59.0% vs 75.4%) and GDP.PDF (35% vs 40% for GPT-5.6 Sol)
  • Non-English production without requalification — model card reports a 5.4-point multilingual safety regression vs 3.7 (lower is better); Google calls it mostly false positives
  • Teams needing price stability beyond 2026 — the introductory rate expires Dec 31, 2026, then doubles to $1.50/$7.50

How it performs by task

Software engineering (DeepSWE v1.1)

Very Good

73.7% Google-reported (74% ±1 independent at high effort) vs 65.3% for 3.7 Flash; outperforms most larger frontier models at a fraction of the cost

Terminal/CLI agent work (Terminal-Bench 2.1)

Very Good

89.4% vs 85.8% for 3.7 (Enterprise guide surface reports 90.8%)

General agentic work (Terminal-Bench 4.0)

Good

19.1% vs 11.2% for 3.7, but still far behind Claude Opus 5's 51.8%

Computer use (OSWorld 2.0)

Good

59.0% vs 50.6% for 3.7; behind Opus 5's 75.4%

Finance and legal agents (Vals Finance Agent V2 / Harvey Legal)

Very Good

61.4% and 10.0% all-pass respectively, both ahead of 3.7; Google says it leads other frontier models

Expert reasoning (HLE-Verified)

Very Good

54.9% vs 53.6% for 3.7

Prompt-injection robustness (Gray Swan)

Very Good

5.5% ASR@15, near Claude Opus 5's 4.8% and well ahead of 3.7 Flash's 9.2%

Pricing

Input

$0.75 / 1M

Output

$3.75 / 1M

Context

1M context, 64K output

View full pricing

Benchmarks

BenchmarkScoreSource
DeepSWE v1.173.7% Source
DeepSWE v1.1 (independent leaderboard, high effort)74% ±1 Source
Terminal-Bench 2.189.4% Source
Terminal-Bench 4.019.1% Source
OSWorld 2.0 (batch tool)59.0% Source
HLE-Verified54.9% Source
Vals Finance Agent v261.4% Source
Gray Swan indirect prompt injection (ASR@15)5.5% Source

No verdict changes yet

The clock starts day one — changes land here as our verdict evolves.

Verification log

No verification checks yet

We haven't logged a verification check for this entry. Once a check runs, its history shows here.

How we evaluate