GPT-Transcribe logo

GPT-Transcribe

OpenAI · Released Jul 2026

Recommended

OpenAI's latest batch speech-to-text model scores 3.31% AA-WER at $0.0045/min — 25% cheaper and measurably more accurate than its gpt-4o-transcribe predecessor. It supports context prompting, keyword hints, and language detection across 22+ languages, making it the best default for asynchronous file transcription on the OpenAI platform. Caveat: no diarization or word-level timestamps; use gpt-4o-transcribe-diarize for speaker labels or gpt-live-transcribe for real-time needs. Best for developers building meeting transcription, media processing, and voice interfaces on the OpenAI API.

Is it right for you?

Good for

  • High-accuracy batch transcription with 3.31% AA-WER across diverse audio conditions
  • Context-aware transcription via free-form prompts, keywords, and language hints for domain-specific audio
  • 22+ language support with detected-language output on completion
  • 25% cheaper than predecessor gpt-4o-transcribe at $0.0045/min

Not good for

  • Multilingual code-switching, noisy audio, or long-form earnings-call audio where specialized models outperform
  • Multi-speaker meetings requiring speaker diarization — use gpt-4o-transcribe-diarize instead
  • Real-time streaming transcription needing sub-second latency — use gpt-live-transcribe

How it performs by task

Batch file transcription

Excellent

Top-tier accuracy at a competitive price; best-in-class on the OpenAI platform for asynchronous workloads

Realtime turn transcription

Very Good

Strong for post-commit turn transcription with language detection, but not designed for live streaming deltas

Multilingual transcription

Very Good

19.27% TER across 22 Common Voice languages — huge leap over Whisper but trails specialized multilingual models on code-switching

Domain-specific audio (medical, legal, technical)

Very Good

Context prompts and keyword hints improve semantic accuracy to 45.2%, but still benefits from domain fine-tuning

Noisy audio transcription

Very Good

Improved over Whisper on real-world noise, but 6.10% WER on Earnings22 shows room for improvement on challenging acoustic conditions

Pricing

Input

$0.0045 / minute

Output

N/A

Context

Per-minute audio billing

View full pricing

Benchmarks

BenchmarkScoreSource
AA-WER v23.31% Source
AA-AgentTalk WER2.22% Source
VoxPopuli-Cleaned-AA WER2.72% Source
Earnings22-Cleaned-AA WER6.10% Source
Common Voice 22-language TER19.27% Source
Context Aware ASR (semantic accuracy)45.2% Source

No verdict changes yet

The clock starts day one — changes land here as our verdict evolves.

Verification log

  • Pricing— No changes

    Automated agent

How we evaluate