Gemini 3.8 Flash TTS and Flash-Lite TTS Hit the Gemini API at $9 and $6 per 1M Audio Tokens, and Both Rates Double on 1 January

Google logoGoogleImportantOctober 3, 2026Models
What happened
Google shipped Gemini 3.8 Flash TTS and Flash-Lite TTS to the Gemini API and AI Studio on 23 September 2026, at $9.00 and $6.00 per 1M audio output tokens.
Why it matters
Both undercut Gemini 3.1 Flash TTS Preview's $20.00 today, but every 3.8 rate doubles on 1 January 2027.
What to do
Test gemini-3.8-flash-tts or gemini-3.8-flash-lite-tts against your current TTS model, and budget at the 2027 rates.

Our verdict: conditional on both. If you generate speech with Gemini 3.1 Flash TTS Preview, test Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS now: both are cheaper per audio token today, and Flash TTS stays cheaper even after the price step. The condition is the calendar. Every 3.8 rate doubles on 1 January 2027, and the quality rankings Google cites are its own reporting of a third-party benchmark, not something we measured. Size any workload at the 2027 rates before you move it.

What happened

Google announced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS on 23 September 2026. Both are available "today" in the Gemini API and Google AI Studio. Flash TTS is also in Gemini Notebook, and Flash-Lite TTS in Google Vids. The Gemini Enterprise API is "coming soon".

  • Coverage: "over 100 languages" and "2,000+ production-ready voices", per Google's launch post.
  • Provenance: every clip is watermarked with SynthID, and Google adds C2PA credentials.
  • Voice replication: Google says you can replicate a voice from "a 30-second audio sample", with built-in consent verification.

The Gemini API pricing page lists these rates per 1M tokens:

Model (API id)Text inputAudio outputFrom 1 Jan 2027 (input / output)
Gemini 3.8 Flash TTS (gemini-3.8-flash-tts)$0.50$9.00$1.00 / $18.00
Gemini 3.8 Flash-Lite TTS (gemini-3.8-flash-lite-tts)$0.50$6.00$1.00 / $12.00
Gemini 3.1 Flash TTS Preview (gemini-3.1-flash-tts-preview)$1.00$20.00no change listed

Batch is half the standard rate on both 3.8 models ($4.50 and $3.00 per 1M audio output tokens through 31 December 2026), and the free tier covers both.

Why it matters

The price gap is the story. Flash TTS audio output costs 55% less than the 3.1 preview today, and Flash-Lite TTS 70% less. After 1 January, Flash TTS ($18.00) is still below the preview's $20.00, and Flash-Lite TTS ($12.00) still 40% below. Text input is half the preview's $1.00 now and matches it from 2027.

That makes the move from the preview an easy call on cost alone. What keeps both models at conditional rather than recommended:

  • The launch rates are temporary. A budget built on $9.00 or $6.00 breaks on 1 January 2027.
  • Quality is Google-reported. Google says Flash TTS took the "#1 overall spot on Hume AI's Voice Design Benchmark (71.4)" and that the two models took the "#1 and #2 spots respectively on Hume AI's Overall Quality Index". Those are Google's citations of a third-party benchmark. We have not run either model.
  • Voice replication is region-locked. "Voice replication through AI Studio is not available in Illinois, Texas, EEA, UK, Switzerland, and India."

What changes for you

  • On the 3.1 preview: run your existing scripts through gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts and compare output quality and cost side by side.
  • Picking a tier: Flash-Lite TTS is the cost tier for high-volume narration and voice-agent cascades; Flash TTS is the one to test for expressive or character voices.
  • Budgeting: size long-running workloads at $18.00 and $12.00 per 1M audio output tokens, not the 2026 rates. Use batch where latency allows to halve either figure.
  • Voice replication: if your users or team sit in Illinois, Texas, the EEA, the UK, Switzerland or India, the AI Studio replication feature is not available to them.

What to do

  1. 1 If you use gemini-3.1-flash-tts-preview, run the same scripts through gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts and compare quality and cost.
  2. 2 Size any long-running workload at the rates from 1 January 2027 ($18.00 and $12.00 per 1M audio output tokens), not the 2026 rates.
  3. 3 If you plan to use voice replication, check the geographic exclusions first: Illinois, Texas, the EEA, the UK, Switzerland and India.

Affected tools & models

Never need to catch up again

The weekly delta — only verdict changes and act-now items. No digest filler.

By subscribing you agree to our Privacy Policy. Unsubscribe anytime.