Gemini 3.8 Flash TTS and Flash-Lite TTS Hit the Gemini API at $9 and $6 per 1M Audio Tokens, and Both Rates Double on 1 January
- What happened
- Google shipped Gemini 3.8 Flash TTS and Flash-Lite TTS to the Gemini API and AI Studio on 23 September 2026, at $9.00 and $6.00 per 1M audio output tokens.
- Why it matters
- Both undercut Gemini 3.1 Flash TTS Preview's $20.00 today, but every 3.8 rate doubles on 1 January 2027.
- What to do
- Test gemini-3.8-flash-tts or gemini-3.8-flash-lite-tts against your current TTS model, and budget at the 2027 rates.
Our verdict: conditional on both. If you generate speech with Gemini 3.1 Flash TTS Preview, test Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS now: both are cheaper per audio token today, and Flash TTS stays cheaper even after the price step. The condition is the calendar. Every 3.8 rate doubles on 1 January 2027, and the quality rankings Google cites are its own reporting of a third-party benchmark, not something we measured. Size any workload at the 2027 rates before you move it.
What happened
Google announced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS on 23 September 2026. Both are available "today" in the Gemini API and Google AI Studio. Flash TTS is also in Gemini Notebook, and Flash-Lite TTS in Google Vids. The Gemini Enterprise API is "coming soon".
- Coverage: "over 100 languages" and "2,000+ production-ready voices", per Google's launch post.
- Provenance: every clip is watermarked with SynthID, and Google adds C2PA credentials.
- Voice replication: Google says you can replicate a voice from "a 30-second audio sample", with built-in consent verification.
The Gemini API pricing page lists these rates per 1M tokens:
| Model (API id) | Text input | Audio output | From 1 Jan 2027 (input / output) |
|---|---|---|---|
Gemini 3.8 Flash TTS (gemini-3.8-flash-tts) | $0.50 | $9.00 | $1.00 / $18.00 |
Gemini 3.8 Flash-Lite TTS (gemini-3.8-flash-lite-tts) | $0.50 | $6.00 | $1.00 / $12.00 |
Gemini 3.1 Flash TTS Preview (gemini-3.1-flash-tts-preview) | $1.00 | $20.00 | no change listed |
Batch is half the standard rate on both 3.8 models ($4.50 and $3.00 per 1M audio output tokens through 31 December 2026), and the free tier covers both.
Why it matters
The price gap is the story. Flash TTS audio output costs 55% less than the 3.1 preview today, and Flash-Lite TTS 70% less. After 1 January, Flash TTS ($18.00) is still below the preview's $20.00, and Flash-Lite TTS ($12.00) still 40% below. Text input is half the preview's $1.00 now and matches it from 2027.
That makes the move from the preview an easy call on cost alone. What keeps both models at conditional rather than recommended:
- The launch rates are temporary. A budget built on $9.00 or $6.00 breaks on 1 January 2027.
- Quality is Google-reported. Google says Flash TTS took the "#1 overall spot on Hume AI's Voice Design Benchmark (71.4)" and that the two models took the "#1 and #2 spots respectively on Hume AI's Overall Quality Index". Those are Google's citations of a third-party benchmark. We have not run either model.
- Voice replication is region-locked. "Voice replication through AI Studio is not available in Illinois, Texas, EEA, UK, Switzerland, and India."
What changes for you
- On the 3.1 preview: run your existing scripts through
gemini-3.8-flash-ttsandgemini-3.8-flash-lite-ttsand compare output quality and cost side by side. - Picking a tier: Flash-Lite TTS is the cost tier for high-volume narration and voice-agent cascades; Flash TTS is the one to test for expressive or character voices.
- Budgeting: size long-running workloads at $18.00 and $12.00 per 1M audio output tokens, not the 2026 rates. Use batch where latency allows to halve either figure.
- Voice replication: if your users or team sit in Illinois, Texas, the EEA, the UK, Switzerland or India, the AI Studio replication feature is not available to them.
What to do
- 1 If you use gemini-3.1-flash-tts-preview, run the same scripts through gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts and compare quality and cost.
- 2 Size any long-running workload at the rates from 1 January 2027 ($18.00 and $12.00 per 1M audio output tokens), not the 2026 rates.
- 3 If you plan to use voice replication, check the geographic exclusions first: Illinois, Texas, the EEA, the UK, Switzerland and India.
Affected tools & models
Never need to catch up again
The weekly delta — only verdict changes and act-now items. No digest filler.