Gemini 3.8 Flash logo

Gemini 3.8 Flash

Google · Veröffentlicht Sept. 2026

Empfohlen

Gemini 3.8 Flash is Google's strongest Flash workhorse yet — a genuine capability upgrade over the recommended 3.7 Flash at the same introductory price, with real gains across software engineering, agents and professional domains. The deciding caveat is Google's own: 3.8 "works harder," so per-task token burn runs well above 3.7, which remains the efficiency-first pick. Adopt 3.8 for long-horizon coding, agents and finance or legal work — benchmark your own workload before the intro price expires.

Ist es das Richtige für dich?

Gut für

  • Long-horizon software engineering — DeepSWE v1.1 73.7% on Google's chart (74% ±1 on the independent leaderboard at high effort), up from 65.3% for 3.7 Flash
  • Autonomous agents and computer use — Terminal-Bench 4.0 19.1% vs 11.2% for 3.7, OSWorld 2.0 59.0% vs 50.6%
  • Knowledge-dense professional domains — HLE-Verified 54.9%; leads 3.7 and other frontier models on Vals Finance Agent V2 (61.4%) and Harvey's Legal Agent Benchmark
  • Google-ecosystem builders — stable GA ID gemini-3.8-flash in Gemini API, AI Studio, Antigravity, Android Studio, Gemini Enterprise; powers Gemini app, AI Mode and Gemini in Sheets

Nicht geeignet für

  • Token-metered, latency-sensitive production — "works harder" means extra reasoning steps and iterative tool calls; Artificial Analysis measured ~40% higher cost per task than 3.7 ($0.40 to $0.58)
  • Hardest frontier tasks — still trails Claude Opus 5 on Terminal-Bench 4.0 (19.1% vs 51.8%), OSWorld 2.0 (59.0% vs 75.4%) and GDP.PDF (35% vs 40% for GPT-5.6 Sol)
  • Non-English production without requalification — model card reports a 5.4-point multilingual safety regression vs 3.7 (lower is better); Google calls it mostly false positives
  • Teams needing price stability beyond 2026 — the introductory rate expires Dec 31, 2026, then doubles to $1.50/$7.50

Leistung nach Aufgabe

Software engineering (DeepSWE v1.1)

Very Good

73.7% Google-reported (74% ±1 independent at high effort) vs 65.3% for 3.7 Flash; outperforms most larger frontier models at a fraction of the cost

Terminal/CLI agent work (Terminal-Bench 2.1)

Very Good

89.4% vs 85.8% for 3.7 (Enterprise guide surface reports 90.8%)

General agentic work (Terminal-Bench 4.0)

Good

19.1% vs 11.2% for 3.7, but still far behind Claude Opus 5's 51.8%

Computer use (OSWorld 2.0)

Good

59.0% vs 50.6% for 3.7; behind Opus 5's 75.4%

Finance and legal agents (Vals Finance Agent V2 / Harvey Legal)

Very Good

61.4% and 10.0% all-pass respectively, both ahead of 3.7; Google says it leads other frontier models

Expert reasoning (HLE-Verified)

Very Good

54.9% vs 53.6% for 3.7

Prompt-injection robustness (Gray Swan)

Very Good

5.5% ASR@15, near Claude Opus 5's 4.8% and well ahead of 3.7 Flash's 9.2%

Preise

Eingabe

$0.75 / 1M

Ausgabe

$3.75 / 1M

Kontext

1M context, 64K output

Alle Preise ansehen

Benchmarks

BenchmarkWertQuelle
DeepSWE v1.173.7% Quelle
DeepSWE v1.1 (independent leaderboard, high effort)74% ±1 Quelle
Terminal-Bench 2.189.4% Quelle
Terminal-Bench 4.019.1% Quelle
OSWorld 2.0 (batch tool)59.0% Quelle
HLE-Verified54.9% Quelle
Vals Finance Agent v261.4% Quelle
Gray Swan indirect prompt injection (ASR@15)5.5% Quelle

Noch keine Urteilsänderungen

Die Uhr läuft ab dem ersten Tag — Änderungen erscheinen hier, sobald sich unser Urteil weiterentwickelt.

Prüfprotokoll

Noch keine Prüfungen

Wir haben für diesen Eintrag noch keine Prüfung erfasst. Sobald eine Prüfung läuft, erscheint ihr Verlauf hier.

Wie wir bewerten