Gemini 3.8 Flash TTS logo

Gemini 3.8 Flash TTS

Google · Released Sep 2026

Conditional

Conditional. Gemini 3.8 Flash TTS is Google's flagship text-to-speech model, available in the Gemini API with voice design, consented voice replication and SynthID watermarking on every clip. The deciding caveat is cost: today's paid rates are launch prices that double when the launch window closes, and its Hume AI leaderboard placements are Google-reported, not measured by us. Suits developers building expressive narration, character voices or dubbing who can budget for the post-launch rate.

Is it right for you?

Good for

  • Expressive narration and character voices; Google's model page calls it a flagship creative model for studio-grade voice fidelity, expressive acting and voice design
  • Voice replication from a 30-second sample with consent verification, SynthID and C2PA credentials; persistent custom voices up to 200 per project with a 1-year TTL
  • Multilingual speech: over 130 languages per Google's model and speech docs
  • Two-speaker dialogue in a single request using prebuilt voices

Not good for

  • Voice replication through Google AI Studio for users in Illinois, Texas, the EEA, the UK, Switzerland or India, where Google does not offer it
  • Dialogue with more than two speakers in one request
  • Anything beyond text in, audio out; the TTS models accept text only and return audio only
  • Budgets built on the launch rate, which is priced through December 31, 2026 and doubles from January 1, 2027

How it performs by task

Expressive narration and voice design

Very Good

Google positions it for expressive acting and voice design and reports top placement on Hume AI's Voice Design Benchmark; vendor-reported, untested by us

Voice cloning

Good

30-second replication with consent checks and watermarking, but Google does not offer it through AI Studio in the EEA, UK, Switzerland, India, Illinois or Texas

Multi-speaker dialogue

Good

Supported, capped at two speakers per request with prebuilt voices

Pricing

Input

$0.50 / 1M (text)

Output

$9.00 / 1M (audio)

Context

Launch rate to 2026-12-31; $1/$18 from 2027

View full pricing

Benchmarks

BenchmarkScoreSource
Hume AI Voice Design Benchmark (Google-reported)71.4 Source
Hume AI Accent Modeling (Google-reported)60.8 Source

No verdict changes yet

The clock starts day one — changes land here as our verdict evolves.

Verification log

  • Pricing— No changes

    Automated agent

How we evaluate