Grok Voice Think Fast 2.0
SpaceXAI · Released Jul 2026
Grok Voice Think Fast 2.0 is SpaceXAI's next-generation speech-to-speech model with a parallel reasoning architecture that makes it substantially smarter than conventional voice models with no latency penalty. Its transcription accuracy destroys dedicated STT competitors in noisy environments — the deciding factor for production voice agents. But at $0.08/min, developer/API-only with no consumer-facing assistant, it's a premium specialist — not for cost-sensitive or general-purpose voice use.
Is it right for you?
Good for
- Parallel reasoning during speech — the model thinks while it speaks, improving intelligence without latency penalty
- Transcription in noisy environments — 10x advantage over dedicated STT models when background noise is present
- Voice agent deployment at scale — 25+ languages, 21 voices, no-code Voice Agent Builder for production
Not good for
- Consumer voice assistant use — model is developer/API-only, no end-user conversational product like ChatGPT Voice
- Cost-sensitive high-volume telephony — at $0.08/min with $0.01/min telephony surcharge, budget STT-only pipelines are cheaper
How it performs by task
Speech-to-speech conversation
SOTA parallel reasoning with 82.9% benchmark score, leading the category
Transcription accuracy
1.5-2x better than dedicated STT models, 10x in noisy settings
Multilingual voice agents
25+ languages with 21 voices, strong but real-world quality varies by language
Cost-effectiveness for high-volume
$0.08/min premium pricing — dedicated STT+TTS pipelines are cheaper at scale
Pricing
Input
$0.08 / min of audio (speech-to-speech)
Output
Included in speech-to-speech rate
Context
API-only
Benchmarks
No verdict changes yet
The clock starts day one — changes land here as our verdict evolves.
Sources
Verification log
- Pricing— No changes
Automated agent