The cheapest MiniMax H3 is here,$0.013/s

Gemini 3.8 Flash TTS

Gemini 3.8 Flash TTS turns scripts into directed performances — voice design, two-speaker dialogue and 130 languages through one Gemini API call.

Gemini 3.8 Flash TTS
Direct expressive speech with Flash TTS or scale cheaply with Flash-Lite TTS
AI Video Prompt Generator

Support

Pro AI Tools

Explore elite tools including the Suno AI Music Generator.

placeholder hero

Why Gemini 3.8 Flash TTS Changes Voice Production

Google's September 23, 2026 release splits Gemini TTS into a creative flagship and a high-throughput workhorse for bulk audio.

  • One Launch, Two Production Targets
    Gemini 3.8 Flash TTS is the creative flagship, while Flash-Lite TTS is the cheaper workhorse for high-volume speech.
  • Directing Delivery, Not Picking a Preset
    Turn-level style instructions, structured speech metadata and inline vocal events set tone, pacing, emotion and accent.
  • Voice Design and Consent-Based Replication
    Describe a voice in natural language, or replicate a real speaker from a reference clip plus a matching consent recording.

Prompting Gemini 3.8 Flash TTS Correctly

Four moves that keep the transcript clean and the performance metadata in charge.

Gemini 3.8 Flash TTS Capabilities in Detail

From performance direction to 130-language coverage, here is what the flagship Gemini TTS model actually does.

Expressive Performance Direction

Turn-level style plus inline laughs, sighs, coughs, breaths and pauses — closer to directing a voice actor than choosing a preset.

Natural-Language Voice Design

Prompts describe age range, personality, accent, vocal texture and role, backed by 2,000+ production voices via the Voices endpoint.

Consent-Gated Voice Replication

A clean reference recording plus a matching consent recording from the same adult speaker, with SynthID watermarking and C2PA credentials.

Two-Speaker Dialogue Staging

Scripts define the conversation for podcasts, educational dialogues, product demos and game scenes without manual line stitching.

Long-Form Stability Without Voice Drift

Google documents stable voice identity, timbre, volume and room tone across multi-minute narration and extended dialogue.

130 Languages and Regional Accents

Flash TTS covers 130 languages against Flash-Lite's 101, with regional accents, minority dialects and IPA pronunciation overrides.

FAQ

Gemini 3.8 Flash TTS — Common Questions

Short answers on Gemini 3.8 Flash TTS pricing, model choice, benchmarks and safeguards.

1

How much does Gemini 3.8 Flash TTS cost per minute?

About 1.35 cents per audio minute at launch rates of $0.50 input and $9 output per million tokens.

2

Should I use Flash TTS or Flash-Lite TTS?

Flash TTS for acting nuance and long-form audio; Flash-Lite TTS for bulk output and low latency.

3

Where does it rank against other voice models?

Google reports 71.4 on Hume's Voice Design Benchmark; Voice Arena places it second at 1,260 Elo.

4

How does it compare with Gemini 3.1 Flash TTS Preview?

Flash-Lite TTS replaces the 3.1 preview, cutting audio output from $20 to $6 per million tokens.

5

Is voice cloning allowed, and what safeguards apply?

Replication needs a reference clip plus a matching consent recording from the same adult speaker.

6

Why are my stage directions being read out loud?

The input is read as a verbatim transcript, so move sustained directions into speech metadata.

Put Gemini 3.8 Flash TTS on Your Own Scripts

Test both models in the Gemini API or Google AI Studio — switching between Gemini 3.8 Flash TTS and Flash-Lite TTS is only a model identifier change. Compare batch and priority inference before you commit a production budget.