Gemini 3.1 Pro logo

Gemini 3.1 Pro

Google · Released Feb 2026

Conditional

Gemini 3.1 Pro is Google's most capable publicly available Pro-tier model, delivering strong reasoning with a 1M-token context window at competitive pricing — and it's the migration target for GitHub Copilot users leaving deprecated Gemini 2.5 Pro today. But it's decidedly previous-gen: Gemini 3.6 and 3.5 Flash now beat it on coding and agentic benchmarks at lower cost. Stick with it if you're deep in the Google Cloud / Vertex AI ecosystem; otherwise, the Flash line or competitors deliver more for less.

Is it right for you?

Good for

  • Deep reasoning on complex scientific and mathematical problems — ARC-AGI-2 77.1% more than doubles Gemini 3 Pro's 31.1%, GPQA Diamond 94.3% leads at launch
  • Very large context analysis (1M tokens) — entire codebases, multi-document research, lengthy contracts in a single session
  • Cost-conscious teams needing Pro-tier reasoning at $2/$12 — 2.5x cheaper input than Claude Opus 4.6 ($5) and 7.5x cheaper than Fable 5 ($15)
  • Google Cloud / Vertex AI ecosystem users who need the strongest available Google model for complex reasoning workloads

Not good for

  • Cutting-edge coding and agentic workflows — Gemini 3.6 Flash and 3.5 Flash surpass it on Terminal-Bench, SWE-Bench Pro, and MCP Atlas at lower cost
  • Security-sensitive production deployments — FAR.AI found 249 universal jailbreaks for ~$278 in July 2026 testing; Claude Fable 5 and GPT-5.6 Sol were unbreachable

How it performs by task

Complex reasoning / science

Excellent

ARC-AGI-2 77.1% (2.5x Gemini 3 Pro), GPQA Diamond 94.3% — top-tier reasoning at launch

Long-context analysis

Excellent

1M token context window, MRCR v2 84.9% at 128K — process entire codebases or multi-document research

Coding

Very Good

SWE-Bench Verified 80.6% is strong but now trails Gemini 3.6 Flash and Claude Opus 5

Agentic workflows

Very Good

APEX-Agents 33.5%, MCP Atlas 69.2% — solid but Flash models are better for production agents

Math

Excellent

LiveCodeBench 2887 Elo, HLE 44.4% no-tools — strong mathematical reasoning

Multimodal understanding

Very Good

MMMU-Pro 80.5% — competitive, natively multimodal across text, audio, images, video

Security / safety

Fair

FAR.AI found 249 universal jailbreaks for ~$278 — significantly weaker safeguards than Claude Fable 5 and GPT-5.6 Sol

Pricing

Input

$2 / 1M

Output

$12 / 1M

Context

1M tokens

View full pricing

Benchmarks

BenchmarkScoreSource
ARC-AGI-277.1% Source
MMLU (MMMLU)92.6% Source
GPQA Diamond94.3% Source
SWE-Bench Verified80.6% Source
SWE-Bench Pro (Public)54.2% Source
Terminal-Bench 2.068.5% Source
LiveCodeBench Pro2887 Elo Source
APEX-Agents33.5% Source
HLE (no tools)44.4% Source
HLE (with tools)51.4% Source
MCP Atlas69.2% Source
BrowseComp85.9% Source
MRCR v2 (128K)84.9% Source
MMMU-Pro80.5% Source
GDPval-AA Elo1317 Source

No verdict changes yet

The clock starts day one — changes land here as our verdict evolves.

Verification log

  • Pricing— No changes

    Automated agent

How we evaluate