Grok 4.5 Ships at $2/$6 — Undercuts Opus by 60%
- What happened
- SpaceXAI released Grok 4.5 on July 8 — its first post-IPO model — at $2/$6 per million tokens and 80 TPS. Trained alongside Cursor IDE data.
- Why it matters
- At $2/$6, Grok 4.5 undercuts Opus 4.8 by 60%+ and GPT-5.6 Sol by 80% — the most aggressive frontier-adjacent pricing from any major lab this year. It resolves SWE Bench Pro with 4.2x fewer output tokens than Opus 4.8.
- What to do
- Cost-sensitive teams running agentic coding loops should evaluate Grok 4.5 now — the pricing is live and the token-efficiency advantage is real. Wait for independent benchmarks before routing mission-critical tasks. EU users: availability expected mid-July.
SpaceXAI released Grok 4.5 on July 8 — its first model since the company went public last month — and the pricing is the real story. At $2 per million input tokens and $6 per million output tokens, Elon Musk's lab is undercutting every frontier-adjacent competitor by a margin wide enough to reset the cost conversation for the entire market.
Our verdict: conditional. Grok 4.5 is in the directory, and for the right buyer — cost-sensitive teams running agentic coding workloads — the value proposition is genuine. The catch: every published benchmark is vendor-run, and Musk's "Opus-class" framing is his own. Independent verification is what separates a pricing move from a procurement recommendation.
Grok 4.5: what SpaceXAI shipped
SpaceXAI published benchmarks Wednesday showing Grok 4.5 competitive with leading models across coding, reasoning, and general knowledge tasks — though just short of best-in-class on most measures (TechCrunch, 2026(opens in new tab)). The company claims "twice greater token efficiency" than other leading models.
The pricing is the headline:
| Model | Input (per 1M) | Output (per 1M) | Grok savings |
|---|---|---|---|
| Grok 4.5 | $2 | $6 | — |
| Claude Opus 4.7 | $5 | $25 | 60% / 76% |
| GPT-5.6 Sol | $5 | $30 | 60% / 80% |
| GPT-5.6 Luna | $1 | $6 | -100% / 0% |
Musk characterized Grok 4.5 as "roughly comparable to Opus 4.7, but much faster" in a follow-up post on X(opens in new tab). "The combination of capability, faster speed and lower cost is what makes it competitive."
SpaceXAI says strong positive feedback from its beta test program led to the accelerated public launch on July 9.
Why it matters
This is the most aggressive pricing from a major AI lab this year. SpaceXAI (formerly xAI) is now a publicly traded company with access to capital markets. A credible frontier-adjacent model at $2/$6 — undercutting the Anthropic flagship by 60% and OpenAI's Sol by 80% — forces every competitor's pricing hand.
The token-efficiency numbers are striking: Grok 4.5 resolves SWE Bench Pro tasks with ~16K output tokens versus Opus 4.8's 67K — a 4.2x advantage. For teams running extensive agentic coding loops where token costs compound, that efficiency gap is real money (SpaceXAI, 2026(opens in new tab)).
But the capability story is nuanced. SpaceXAI's own published benchmarks place Grok 4.5 behind Fable 5 and GPT-5.5 on most coding evaluations, though it leads Opus 4.8 on DeepSWE 1.0, SWE Marathon, and Terminal Bench 2.1. Trained alongside Cursor IDE interaction data, Grok 4.5 is purpose-built for long-running tool use and multi-step engineering — not for pushing the absolute frontier on raw reasoning benchmarks.
The caveats
We haven't tested Grok 4.5 independently. Every benchmark cited is vendor-reported — same caveat we apply to every model launch. Musk's "Opus-class" claim is directionally accurate by his own qualified framing (comparable to Opus 4.7, not the current Opus 4.8), but the model doesn't lead the frontier on any independently verified metric.
The competitive landscape is also a moving target: GPT-5.6 Sol is expected to launch publicly today (July 9) (TechCrunch, 2026(opens in new tab)). Grok 4.5 enters a market where the frontier is being redefined in real time — and it's not yet available in the EU (mid-July expected).
FAQ
Is Grok 4.5 worth switching to right now?
For cost-sensitive agentic coding, yes — $2/$6 pricing with 80 TPS and 4.2x token efficiency over Opus 4.8 is a compelling value proposition. For mission-critical workloads where a single wrong answer is expensive, wait for independent benchmarks.
How does Grok 4.5 actually compare to Opus 4.8?
SpaceXAI's vendor-reported benchmarks show Grok 4.5 trailing Opus 4.8 on most standard coding evals but leading on DeepSWE 1.0, SWE Marathon, and Terminal Bench 2.1. Musk himself qualified it as "roughly comparable to Opus 4.7" — not the current Opus 4.8. Independent verification is pending.
When will independent benchmarks arrive?
No timeline confirmed. Artificial Analysis and other third-party evaluators typically publish within days of a public API launch. We'll update the directory verdict when they land.
What changes for you
If you run cost-sensitive agentic coding: Grok 4.5's $2/$6 pricing — paired with 80 TPS speed — makes it the cheapest model in the frontier-adjacent conversation. For teams that already route high-volume tasks to budget tiers, Grok 4.5 offers near-Opus performance on coding benchmarks at a fraction of the per-token cost.
If you need the absolute best capability: This isn't your model. Fable 5 and Opus 4.8 lead where raw reasoning and frontier coding depth matter. Use Grok 4.5 where token efficiency compounds across long agent loops — not where a single wrong answer is expensive.
If you're in the EU: Wait. Availability is expected mid-July but not confirmed.
Our stance is clear: the pricing is real, the launch is real, and the token-efficiency numbers — if they hold under independent scrutiny — are the strongest argument for Grok 4.5 over every Anthropic and OpenAI alternative. What's missing is third-party verification. We'll update the verdict when independent benchmarks land.
Affected tools & models
Never need to catch up again
The weekly delta — only verdict changes and act-now items. No digest filler.