Gemini 3.1 Pro
Google · Released Feb 2026
Gemini 3.1 Pro is Google's most capable publicly available Pro-tier model, delivering strong reasoning with a 1M-token context window at competitive pricing — and it's the migration target for GitHub Copilot users leaving deprecated Gemini 2.5 Pro today. But it's decidedly previous-gen: Gemini 3.6 and 3.5 Flash now beat it on coding and agentic benchmarks at lower cost. Stick with it if you're deep in the Google Cloud / Vertex AI ecosystem; otherwise, the Flash line or competitors deliver more for less.
Is it right for you?
Good for
- Deep reasoning on complex scientific and mathematical problems — ARC-AGI-2 77.1% more than doubles Gemini 3 Pro's 31.1%, GPQA Diamond 94.3% leads at launch
- Very large context analysis (1M tokens) — entire codebases, multi-document research, lengthy contracts in a single session
- Cost-conscious teams needing Pro-tier reasoning at $2/$12 — 2.5x cheaper input than Claude Opus 4.6 ($5) and 7.5x cheaper than Fable 5 ($15)
- Google Cloud / Vertex AI ecosystem users who need the strongest available Google model for complex reasoning workloads
Not good for
- Cutting-edge coding and agentic workflows — Gemini 3.6 Flash and 3.5 Flash surpass it on Terminal-Bench, SWE-Bench Pro, and MCP Atlas at lower cost
- Security-sensitive production deployments — FAR.AI found 249 universal jailbreaks for ~$278 in July 2026 testing; Claude Fable 5 and GPT-5.6 Sol were unbreachable
How it performs by task
Complex reasoning / science
ARC-AGI-2 77.1% (2.5x Gemini 3 Pro), GPQA Diamond 94.3% — top-tier reasoning at launch
Long-context analysis
1M token context window, MRCR v2 84.9% at 128K — process entire codebases or multi-document research
Coding
SWE-Bench Verified 80.6% is strong but now trails Gemini 3.6 Flash and Claude Opus 5
Agentic workflows
APEX-Agents 33.5%, MCP Atlas 69.2% — solid but Flash models are better for production agents
Math
LiveCodeBench 2887 Elo, HLE 44.4% no-tools — strong mathematical reasoning
Multimodal understanding
MMMU-Pro 80.5% — competitive, natively multimodal across text, audio, images, video
Security / safety
FAR.AI found 249 universal jailbreaks for ~$278 — significantly weaker safeguards than Claude Fable 5 and GPT-5.6 Sol
Pricing
Input
$2 / 1M
Output
$12 / 1M
Context
1M tokens
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| ARC-AGI-2 | 77.1% | Source |
| MMLU (MMMLU) | 92.6% | Source |
| GPQA Diamond | 94.3% | Source |
| SWE-Bench Verified | 80.6% | Source |
| SWE-Bench Pro (Public) | 54.2% | Source |
| Terminal-Bench 2.0 | 68.5% | Source |
| LiveCodeBench Pro | 2887 Elo | Source |
| APEX-Agents | 33.5% | Source |
| HLE (no tools) | 44.4% | Source |
| HLE (with tools) | 51.4% | Source |
| MCP Atlas | 69.2% | Source |
| BrowseComp | 85.9% | Source |
| MRCR v2 (128K) | 84.9% | Source |
| MMMU-Pro | 80.5% | Source |
| GDPval-AA Elo | 1317 | Source |
No verdict changes yet
The clock starts day one — changes land here as our verdict evolves.
Sources
- Google DeepMind: Gemini 3.1 Pro Model CardJul 2026
- FAR.AI AI Security LeaderboardJul 2026
- Google AI for Developers — Official pricing page (no 3.5 Pro yet)Jun 2026
- Google Blog: Gemini 3.1 Pro announcementJul 2026
- Metacto: Gemini API PricingJul 2026
- nxcode.io: Gemini 3.1 Pro Complete GuideJul 2026
- GitHub Changelog: Gemini 3.1 Pro in CopilotJul 2026
- GitHub Changelog: Gemini 2.5 Pro deprecation Jul 31, 2026Jul 2026
Verification log
- Pricing— No changes
Automated agent