G

Grok 4.5

SpaceXAI · Released Jul 2026

Conditional

Competitively priced coding/agentic model at $2/$6 per 1M tokens with 500K context — xAI's published benchmarks place it mid-pack behind Fable 5 and GPT-5.5, though it leads Opus 4.8 on DeepSWE and Terminal-Bench 2.1. The real story is cost efficiency: 4.2x more token-efficient than Opus 4.8 on SWE-Bench Pro. Best for teams needing strong coding at a fraction of frontier pricing. Not yet available in the EU.

Is it right for you?

Good for

  • Cost-sensitive coding workloads — $2/$6 pricing undercuts Opus by 60% on input, 76% on output
  • Token-efficient agentic tasks — resolves SWE Bench Pro with ~16K output tokens, 4.2x fewer than Opus 4.8
  • Long-running tool-use and multi-step engineering — trained alongside Cursor for agentic coding workflows
  • Office productivity — built-in Excel, PowerPoint, Word integration via Grok Build

Not good for

  • Teams needing the absolute frontier — Fable 5 and GPT-5.5 lead on most coding benchmarks
  • EU-based projects — not yet available in EU regions (mid-July 2026 expected)
  • Workloads requiring independent benchmark verification — all current scores are vendor-reported

How it performs by task

Code generation

Very Good

Strong on SWE Bench Pro (64.7%) and Terminal Bench 2.1 (83.3%); mid-pack among frontier models

Agentic coding

Very Good

Built for long-running tool use with Cursor training data; 4.2x token efficiency advantage

Knowledge work

Good

Solid but trails Opus 4.8 and Fable 5 on knowledge benchmarks per vendor charts

Reasoning

Good

Configurable reasoning_effort (low/medium/high); adequate but not class-leading

Multimodal

Good

Text + image input, text output only; no image generation natively

Pricing

Input

$2 / 1M tokens

Output

$6 / 1M tokens

Context

500K tokens

View full pricing

Benchmarks

BenchmarkScoreSource
DeepSWE 1.062.0% Source
DeepSWE 1.153% Source
SWE Marathon29.0% Source
Terminal Bench 2.183.3% Source
SWE Bench Pro64.7% Source

Verdict history

Jul 15, 2026
copy edit to brevity standard — detail preserved in existing structured children

Verification log

  • Pricing— No changes

    Automated agent

  • Pricing— No changes

    Automated agent

  • Profile— No changes

    Imported at launch

  • Pricing— No changes

    Imported at launch

How we evaluate