Gemini 3.8 Flash TTS
Gemini 3.8 Flash TTS turns scripts into directed performances — voice design, two-speaker dialogue and 130 languages through one Gemini API call.
Support
Pro AI Tools
Explore elite tools including the Suno AI Music Generator.
MiniMax H3
MiniMax H3 AI Video Generator
Seedance 2.5
The Future of AI Video Is Here.

Seedance 2.0
The Future of AI Video Is Here.

Veo3.1
Create Stunning Videos with Veo3.1
FLUX 3 Video Generator

Kling 3.0
Next-Gen AI Video Generator
Grok Video Generator
Create Videos from Text or Images with AI
MiniMax H3 video generator
MiniMax H3 AI Video Generator

Why Gemini 3.8 Flash TTS Changes Voice Production
Google's September 23, 2026 release splits Gemini TTS into a creative flagship and a high-throughput workhorse for bulk audio.
- One Launch, Two Production TargetsGemini 3.8 Flash TTS is the creative flagship, while Flash-Lite TTS is the cheaper workhorse for high-volume speech.
- Directing Delivery, Not Picking a PresetTurn-level style instructions, structured speech metadata and inline vocal events set tone, pacing, emotion and accent.
- Voice Design and Consent-Based ReplicationDescribe a voice in natural language, or replicate a real speaker from a reference clip plus a matching consent recording.
Prompting Gemini 3.8 Flash TTS Correctly
Four moves that keep the transcript clean and the performance metadata in charge.
Gemini 3.8 Flash TTS Capabilities in Detail
From performance direction to 130-language coverage, here is what the flagship Gemini TTS model actually does.
Expressive Performance Direction
Turn-level style plus inline laughs, sighs, coughs, breaths and pauses — closer to directing a voice actor than choosing a preset.
Natural-Language Voice Design
Prompts describe age range, personality, accent, vocal texture and role, backed by 2,000+ production voices via the Voices endpoint.
Consent-Gated Voice Replication
A clean reference recording plus a matching consent recording from the same adult speaker, with SynthID watermarking and C2PA credentials.
Two-Speaker Dialogue Staging
Scripts define the conversation for podcasts, educational dialogues, product demos and game scenes without manual line stitching.
Long-Form Stability Without Voice Drift
Google documents stable voice identity, timbre, volume and room tone across multi-minute narration and extended dialogue.
130 Languages and Regional Accents
Flash TTS covers 130 languages against Flash-Lite's 101, with regional accents, minority dialects and IPA pronunciation overrides.
Gemini 3.8 Flash TTS — Common Questions
Short answers on Gemini 3.8 Flash TTS pricing, model choice, benchmarks and safeguards.
How much does Gemini 3.8 Flash TTS cost per minute?
About 1.35 cents per audio minute at launch rates of $0.50 input and $9 output per million tokens.
Should I use Flash TTS or Flash-Lite TTS?
Flash TTS for acting nuance and long-form audio; Flash-Lite TTS for bulk output and low latency.
Where does it rank against other voice models?
Google reports 71.4 on Hume's Voice Design Benchmark; Voice Arena places it second at 1,260 Elo.
How does it compare with Gemini 3.1 Flash TTS Preview?
Flash-Lite TTS replaces the 3.1 preview, cutting audio output from $20 to $6 per million tokens.
Is voice cloning allowed, and what safeguards apply?
Replication needs a reference clip plus a matching consent recording from the same adult speaker.
Why are my stage directions being read out loud?
The input is read as a verbatim transcript, so move sustained directions into speech metadata.
Put Gemini 3.8 Flash TTS on Your Own Scripts
Test both models in the Gemini API or Google AI Studio — switching between Gemini 3.8 Flash TTS and Flash-Lite TTS is only a model identifier change. Compare batch and priority inference before you commit a production budget.
