Gemini 3.8 Flash logo

Gemini 3.8 Flash

Google · Sorti sept. 2026

Recommandé

Gemini 3.8 Flash is Google's strongest Flash workhorse yet — a genuine capability upgrade over the recommended 3.7 Flash at the same introductory price, with real gains across software engineering, agents and professional domains. The deciding caveat is Google's own: 3.8 "works harder," so per-task token burn runs well above 3.7, which remains the efficiency-first pick. Adopt 3.8 for long-horizon coding, agents and finance or legal work — benchmark your own workload before the intro price expires.

Est-ce fait pour vous ?

Recommandé pour

  • Long-horizon software engineering — DeepSWE v1.1 73.7% on Google's chart (74% ±1 on the independent leaderboard at high effort), up from 65.3% for 3.7 Flash
  • Autonomous agents and computer use — Terminal-Bench 4.0 19.1% vs 11.2% for 3.7, OSWorld 2.0 59.0% vs 50.6%
  • Knowledge-dense professional domains — HLE-Verified 54.9%; leads 3.7 and other frontier models on Vals Finance Agent V2 (61.4%) and Harvey's Legal Agent Benchmark
  • Google-ecosystem builders — stable GA ID gemini-3.8-flash in Gemini API, AI Studio, Antigravity, Android Studio, Gemini Enterprise; powers Gemini app, AI Mode and Gemini in Sheets

Déconseillé pour

  • Token-metered, latency-sensitive production — "works harder" means extra reasoning steps and iterative tool calls; Artificial Analysis measured ~40% higher cost per task than 3.7 ($0.40 to $0.58)
  • Hardest frontier tasks — still trails Claude Opus 5 on Terminal-Bench 4.0 (19.1% vs 51.8%), OSWorld 2.0 (59.0% vs 75.4%) and GDP.PDF (35% vs 40% for GPT-5.6 Sol)
  • Non-English production without requalification — model card reports a 5.4-point multilingual safety regression vs 3.7 (lower is better); Google calls it mostly false positives
  • Teams needing price stability beyond 2026 — the introductory rate expires Dec 31, 2026, then doubles to $1.50/$7.50

Performances par tâche

Software engineering (DeepSWE v1.1)

Very Good

73.7% Google-reported (74% ±1 independent at high effort) vs 65.3% for 3.7 Flash; outperforms most larger frontier models at a fraction of the cost

Terminal/CLI agent work (Terminal-Bench 2.1)

Very Good

89.4% vs 85.8% for 3.7 (Enterprise guide surface reports 90.8%)

General agentic work (Terminal-Bench 4.0)

Good

19.1% vs 11.2% for 3.7, but still far behind Claude Opus 5's 51.8%

Computer use (OSWorld 2.0)

Good

59.0% vs 50.6% for 3.7; behind Opus 5's 75.4%

Finance and legal agents (Vals Finance Agent V2 / Harvey Legal)

Very Good

61.4% and 10.0% all-pass respectively, both ahead of 3.7; Google says it leads other frontier models

Expert reasoning (HLE-Verified)

Very Good

54.9% vs 53.6% for 3.7

Prompt-injection robustness (Gray Swan)

Very Good

5.5% ASR@15, near Claude Opus 5's 4.8% and well ahead of 3.7 Flash's 9.2%

Tarifs

Entrée

$0.75 / 1M

Sortie

$3.75 / 1M

Contexte

1M context, 64K output

Voir tous les tarifs

Tests de performance

TestScoreSource
DeepSWE v1.173.7% Source
DeepSWE v1.1 (independent leaderboard, high effort)74% ±1 Source
Terminal-Bench 2.189.4% Source
Terminal-Bench 4.019.1% Source
OSWorld 2.0 (batch tool)59.0% Source
HLE-Verified54.9% Source
Vals Finance Agent v261.4% Source
Gray Swan indirect prompt injection (ASR@15)5.5% Source

Aucun changement de verdict pour l’instant

L’horloge tourne dès le premier jour : les changements apparaîtront ici à mesure que notre verdict évolue.

Journal de vérification

Aucune vérification pour le moment

Nous n’avons pas encore enregistré de vérification pour cette entrée. Dès qu’une vérification s’exécute, son historique apparaît ici.

Comment nous évaluons