Grok Voice Think Fast 2.0 Ships — 82.9% Score at $0.08/Min

SpaceXAI logoSpaceXAIFYIJuly 30, 2026Product Launches
What happened
SpaceXAI released Grok Voice Think Fast 2.0 — 82.9% AA Quality Index, 0.70s time-to-first-audio, 1.5–2.0× transcription gains, $0.08/min pricing. Grok-Voice-Latest auto-migrates Aug 5.
Why it matters
Voice AI is iterating faster than text models did at the same stage — and there's no clear winner. Grok Voice 2.0 at API-accessible $0.08/min pressures OpenAI's ChatGPT-only voice strategy.
What to do
Test Grok Voice 2.0 against your voice agent use case before the Aug 5 cutover. Build model-swappable voice architecture — the leader in August won't be the leader in October.

SpaceXAI's Grok Voice Think Fast 2.0 is the new speech-to-speech benchmark leader — 82.9% on Artificial Analysis' Quality Index, 0.70s time-to-first-audio, and a 7.2-point gain over v1.0. At $0.08/minute, it's competitively priced against a voice model field that's iterating faster than text models did at the same stage. The August 5 auto-migration of grok-voice-latest means every Grok Voice API user gets the upgrade with zero config changes — but it also means if you haven't tested v2.0 yet, you have six days.

What Happened

SpaceXAI announced Grok Voice Think Fast 2.0 on July 29 via the x.ai blog(opens in new tab). The speech-to-speech model delivers across four dimensions:

Benchmark leadership. On Artificial Analysis' Speech-to-Speech Quality Index, Grok Voice Think Fast 2.0 scores 82.9% — ahead of GPT-Realtime-2.1 (79.1%) and Gemini 3.1 Flash (69.5%). Specific highlights:

MetricGrok Voice 2.0Grok Voice 1.0GPT-Realtime-2.1Gemini 3.1 Flash
AA Quality Index82.9%75.7%79.1%69.5%
Conversational Dynamics95.1%77.8%95.7%74.3%
Agentic Performance (τ-voice)56.5%52.1%45.7%37.7%
Time to First Audio0.70s1.25s2.98s

Transcription accuracy. Outperforms dedicated transcription models Deepgram Nova 3 and ElevenLabs Scribe v2 by 1.5–2.0× across 24 languages. The gap widens to ~10× in noisy, real-world settings with telephony compression — this is a speech-to-speech model that handles transcription better than purpose-built speech-to-text tools.

Reasoning efficiency. Grok Voice Think Fast 2.0 uses 0.4× the reasoning tokens of v1.0 while reasoning in parallel with speech. Tool calls execute before the end of the agent's first sentence — snappier, cheaper, same intelligence.

August 5 auto-migration. The grok-voice-latest endpoint transitions from grok-voice-think-fast-1.0 to grok-voice-think-fast-2.0 on August 5, 2026. No action needed to upgrade. To stay on v1.0, pin grok-voice-think-fast-1.0 before then.

$0.08/minute pricing. Enterprise API pricing with no subscription gate — available now via the x.ai console.

Why It Matters

The voice model category is where text models were in early 2025: rapid iteration, aggressive pricing, and no entrenched winner. Grok Voice 2.0 at $0.08/minute with API-accessible full-duplex capability is a direct competitive shot across the bow of OpenAI's voice models — which currently gate real-time voice behind ChatGPT subscriptions with no standalone API pricing for full-duplex.

For developers building voice agents, meeting assistants, or real-time conversational AI: the model landscape is fragmenting fast. You now have GPT-Live (ChatGPT-only, no API), GPT-Transcribe (batch, API), GPT-Live-Transcribe (streaming, API), and Grok Voice 2.0 (full-duplex, API, $0.08/min). Each has different latency, pricing, and API-access tradeoffs. The right choice depends heavily on your use case — and the landscape will look different by October.

SpaceXAI is also using Grok Voice 2.0 in production on Starlink's customer support line (+1 888 GO STARLINK), where A/B testing showed significant gains in sales conversion and support containment rates. That's a real-world validation signal — not just benchmarks.

What Changes for You

  1. Test v2.0 before August 5. If you're on grok-voice-latest, the cutover is automatic. v2.0's faster response time and shorter sentences change conversation pacing — test your integration before the endpoint switches.
  2. Benchmark against your voice use case, not just leaderboards. The 0.70s time-to-first-audio and 56.5% agentic performance score matter more for real-time voice agents than the overall quality index. Run your own latency and task-completion tests.
  3. Don't lock into a single voice model yet. The category is evolving faster than text models did at the same stage. Build your voice pipeline with model-swappable architecture — the leader in August won't necessarily be the leader in October.

Affected tools & models

Never need to catch up again

The weekly delta — only verdict changes and act-now items. No digest filler.

By subscribing you agree to our Privacy Policy. Unsubscribe anytime.