GPT-Transcribe
OpenAI · Released Jul 2026
OpenAI's latest batch speech-to-text model scores 3.31% AA-WER at $0.0045/min — 25% cheaper and measurably more accurate than its gpt-4o-transcribe predecessor. It supports context prompting, keyword hints, and language detection across 22+ languages, making it the best default for asynchronous file transcription on the OpenAI platform. Caveat: no diarization or word-level timestamps; use gpt-4o-transcribe-diarize for speaker labels or gpt-live-transcribe for real-time needs. Best for developers building meeting transcription, media processing, and voice interfaces on the OpenAI API.
Is it right for you?
Good for
- High-accuracy batch transcription with 3.31% AA-WER across diverse audio conditions
- Context-aware transcription via free-form prompts, keywords, and language hints for domain-specific audio
- 22+ language support with detected-language output on completion
- 25% cheaper than predecessor gpt-4o-transcribe at $0.0045/min
Not good for
- Multilingual code-switching, noisy audio, or long-form earnings-call audio where specialized models outperform
- Multi-speaker meetings requiring speaker diarization — use gpt-4o-transcribe-diarize instead
- Real-time streaming transcription needing sub-second latency — use gpt-live-transcribe
How it performs by task
Batch file transcription
Top-tier accuracy at a competitive price; best-in-class on the OpenAI platform for asynchronous workloads
Realtime turn transcription
Strong for post-commit turn transcription with language detection, but not designed for live streaming deltas
Multilingual transcription
19.27% TER across 22 Common Voice languages — huge leap over Whisper but trails specialized multilingual models on code-switching
Domain-specific audio (medical, legal, technical)
Context prompts and keyword hints improve semantic accuracy to 45.2%, but still benefits from domain fine-tuning
Noisy audio transcription
Improved over Whisper on real-world noise, but 6.10% WER on Earnings22 shows room for improvement on challenging acoustic conditions
Pricing
Input
$0.0045 / minute
Output
N/A
Context
Per-minute audio billing
Benchmarks
No verdict changes yet
The clock starts day one — changes land here as our verdict evolves.
Sources
- OpenAI API — Realtime Transcription GuideJul 2026
- OpenAI Developers on X: GPT-Transcribe + GPT-Live-Transcribe announcementJul 2026
- Artificial Analysis — Speech-to-Text LeaderboardJul 2026
- Artificial Analysis on X: GPT Transcribe benchmarksJul 2026
- GIGAZINE — GPT Transcribe / GPT Live Transcribe Release CoverageJul 2026
- OpenAI Developer Pricing DocsJun 2026
- OpenAI API — GPT Transcribe Model PageJul 2026
Verification log
- Pricing— No changes
Automated agent