Gemini 3.1 Pro logo

Gemini 3.1 Pro

Google · Veröffentlicht Feb. 2026

Bedingt

Googles leistungsfähigstes öffentlich verfügbares Pro-Tier-Modell mit 1M-Token-Kontext zu wettbewerbsfähigen Preisen. Vorgänger-Generation: Gemini 3.6 und 3.5 Flash schlagen es jetzt bei Coding- und agentischen Benchmarks zu niedrigeren Kosten. Bleib dabei in Google Cloud/Vertex AI; ansonsten liefern die Flash-Reihe oder Wettbewerber mehr.

Ist es das Richtige für dich?

Gut für

  • Tiefgehendes Reasoning bei komplexen wissenschaftlichen und mathematischen Problemen
  • Analyse sehr großer Kontexte mit 1M Tokens
  • Kostenbewusste Teams, die Pro-Tier-Reasoning für $2/$12 benötigen
  • Nutzer des Google Cloud/Vertex AI-Ökosystems

Nicht geeignet für

  • Modernste Coding- und agentische Workflows — Flash-Modelle übertreffen es
  • Sicherheitssensible Deployments — FAR.AI fand 249 universelle Jailbreaks für ~$278

Leistung nach Aufgabe

Complex reasoning / science

Excellent

ARC-AGI-2 77.1% (2.5x Gemini 3 Pro), GPQA Diamond 94.3% — top-tier reasoning at launch

Long-context analysis

Excellent

1M token context window, MRCR v2 84.9% at 128K — process entire codebases or multi-document research

Coding

Very Good

SWE-Bench Verified 80.6% is strong but now trails Gemini 3.6 Flash and Claude Opus 5

Agentic workflows

Very Good

APEX-Agents 33.5%, MCP Atlas 69.2% — solid but Flash models are better for production agents

Math

Excellent

LiveCodeBench 2887 Elo, HLE 44.4% no-tools — strong mathematical reasoning

Multimodal understanding

Very Good

MMMU-Pro 80.5% — competitive, natively multimodal across text, audio, images, video

Security / safety

Fair

FAR.AI found 249 universal jailbreaks for ~$278 — significantly weaker safeguards than Claude Fable 5 and GPT-5.6 Sol

Preise

Eingabe

$2 / 1M

Ausgabe

$12 / 1M

Kontext

1M tokens

Alle Preise ansehen

Benchmarks

BenchmarkWertQuelle
ARC-AGI-277.1% Quelle
MMLU (MMMLU)92.6% Quelle
GPQA Diamond94.3% Quelle
SWE-Bench Verified80.6% Quelle
SWE-Bench Pro (Public)54.2% Quelle
Terminal-Bench 2.068.5% Quelle
LiveCodeBench Pro2887 Elo Quelle
APEX-Agents33.5% Quelle
HLE (no tools)44.4% Quelle
HLE (with tools)51.4% Quelle
MCP Atlas69.2% Quelle
BrowseComp85.9% Quelle
MRCR v2 (128K)84.9% Quelle
MMMU-Pro80.5% Quelle
GDPval-AA Elo1317 Quelle

Noch keine Urteilsänderungen

Die Uhr läuft ab dem ersten Tag — Änderungen erscheinen hier, sobald sich unser Urteil weiterentwickelt.

Prüfprotokoll

  • Preise— Keine Änderungen

    Automatisierter Agent

Wie wir bewerten