Grok 4.6 Ships at $2/$6 — 61 AA Index, Ties GPT-5.6 Sol
- What happened
- SpaceXAI launched Grok 4.6 on Aug 12 at $2/$6, scoring 61 on the AA Intelligence Index — tying GPT-5.6 Sol for third, behind Claude Opus 5 and Claude Fable 5.
- Why it matters
- It's frontier-competitive and the cheapest in its class — the model to watch for long-running agents and visual/interactive work.
- What to do
- Cursor and Grok Build users should try it now (2x included usage this week); everyone else, hold for broader benchmarks before betting production on it.
Grok 4.6 shipped on August 12 at the same $2 / $6 per 1M tokens as Grok 4.5 — and the first independent read is strong: 61 on the Artificial Analysis Intelligence Index, a nine-benchmark composite that ties GPT-5.6 Sol and trails Claude Opus 5 (63) and Claude Fable 5 (62). We're moving the directory verdict from pending → conditional — not because the benchmarks are weak, but because a model that shipped yesterday hasn't earned recommended yet.
What happened
SpaceXAI launched Grok 4.6 with pricing unchanged from Grok 4.5, plus a "Fast variant" at 2x price for higher throughput. The headline numbers (Artificial Analysis, 2026):
| Metric | Grok 4.6 | GPT-5.6 Sol | Fable 5 Max | Grok 4.5 |
|---|---|---|---|---|
| AA Intelligence Index | 61 | 61 | 62 | 56 |
| GDPVal-AA v2 | 1,753 | 1,728 | 1,741 | 1,526 |
| CursorBench v3.2 | 69.9% | 67.2% | 70.5% | 66.7% |
The table above compares a four-model subset — Claude Opus 5 (63) sits above Grok 4.6 on the AA Intelligence Index composite.
Three things stand out:
- GDPVal-AA 1,753 — behind only Claude Opus 5, ahead of both Fable 5 Max and GPT-5.6 Sol.
- A 5-point composite jump over Grok 4.5 (56 → 61) at the same price.
- 2x included usage in Grok Build and Cursor for the first week — SpaceXAI is paying for the evaluation window.
It ships today across Cursor, Grok Build, the API, OpenRouter, Vercel, and Cloudflare. The launch post frames it for long-running agents and more ambitious interactive/visual work: staying on task across many steps, and stronger first-pass app structure and visual language (xAI, 2026).
Why it matters
Grok 4.6 is frontier-competitive at the cheapest price in its class. Tying GPT-5.6 Sol on the composite — at $2/$6, unchanged from Grok 4.5 — resets the cost-per-capability math for agent workloads. If the GDPVal-AA lead holds, it becomes the default reach for teams that burn tokens on coding agents.
But independent verification is exactly one evaluator deep. Artificial Analysis is the only third-party read so far. A model that shipped yesterday has no long-horizon reliability record — no multi-step agent stress tests, no multi-day task-completion data. That's the gap between conditional and recommended, and a day-one composite score doesn't close it.
The positioning is agent-first. The launch leads with persistence across many steps and first-pass structure and visuals, aimed at Cursor and Grok Build users — not a general-purpose assistant play.
What changes for you
On Cursor or Grok Build: try it this week. The 2x included usage makes it a near-free evaluation, and Grok 4.5 users get a meaningful upgrade at the same price.
Betting production: hold. Let the broader benchmark set land before routing critical workloads. One composite score, however strong, is not a reliability record.
Need throughput: watch the Fast variant. It runs at 2x price; worth it only if your workload is latency-bound, not token-bound.
FAQ
What changed between Grok 4.6 and Grok 4.5? Pricing is unchanged at $2/$6, but Grok 4.6 adds a 5-point composite jump (56 → 61), a GDPVal-AA result behind only Claude Opus 5, and — per the launch post — better persistence on long-running agent tasks plus stronger first-pass app structure and visual language (xAI, 2026).
Is Grok 4.6 better than GPT-5.6 Sol? They tie at 61 on the AA Intelligence Index. Grok 4.6 leads GPT-5.6 Sol on GDPVal-AA (1,753 vs 1,728) and CursorBench (69.9% vs 67.2%), while Claude Opus 5 (63) and Claude Fable 5 (62) still sit ahead on the composite. At $2/$6, Grok 4.6 is the cheaper route to that capability tier — with far less of a track record.
Should I move production workloads to Grok 4.6 now? Not yet. It's one evaluator deep and one day old. Evaluate this week — especially on Cursor and Grok Build with the 2x included usage — and let independent agentic and long-horizon benchmarks land before it touches production.
Launch completes the pending entity: 61 AA Intelligence Index (ties GPT-5.6 Sol, behind Claude Opus 5 and Claude Fable 5) at $2/$6, with GDPVal-AA behind only Claude Opus 5.
What to do
- 1 If you're on Cursor or Grok Build, try Grok 4.6 this week — the 2x included usage makes it a low-cost evaluation.
- 2 Hold production bets until the broader benchmark set lands; one evaluator (Artificial Analysis) is the only independent read so far.
- 3 Watch the 'Fast variant' at 2x price only if your workload is latency-bound, not token-bound.
Affected tools & models
Never need to catch up again
The weekly delta — only verdict changes and act-now items. No digest filler.