Gemini 3.8 Flash TTS and Flash-Lite TTS add custom voices and script direction

Google has released two speech generation models: Flash TTS for creative voice direction and Flash-Lite TTS for high-volume audio workflows.
AI Neural Narration
48kHz StudioFish Audio Neural Engine · Natural editorial narration
Key Takeaways
- check_circleFlash TTS and Flash-Lite TTS serve different production needs.
- check_circleEvaluate a voice on the actual script and language your audience hears.
- check_circleKeep approval records for voices used in customer-facing content.
Two new speech generation options
Google announced both models on 23 September 2026. Flash TTS focuses on voice design and performance direction; Flash-Lite TTS targets efficient volume production. Both support multilingual audio.
Developer rollout is through the Gemini API and AI Studio; enterprise API access is coming soon. Consumer surfaces include Gemini Notebook and Google Vids. Replication requires consent verification and has regional restrictions. Remixing is a future feature.
Choose for the finished recording
Our assessment: an audio team should compare the models using a finished script, rather than a few flattering sample sentences. Include abbreviations, numbers, customer names and words that are easy to mispronounce. Listen for changes in character or delivery across a longer recording. Ask a listener who knows the target language to assess the result.
For a support agent, clarity and response time may matter more than dramatic range. For a narrated lesson, pacing and consistent pronunciation may determine whether the audio is useful. Write down the acceptance criteria before selecting a voice, then score both models against those criteria. Keep the reference script stable so revisions can be compared fairly.
Make the production workflow reviewable
Save the approved script, voice settings and final recording together. That gives an editor a way to reproduce a clip when a price, date or product name changes. Record which version was actually published. A regenerated sentence should be checked in context, because a correct standalone clip can still sound out of place in the complete recording.
For large dubbing jobs, calculate the cost per approved minute after regeneration and review. Build a pronunciation list for recurring terms and sample each language before expanding the batch. Where a voice is based on a person, document the permission and permitted uses before recording. These are editorial recommendations, not a claim that a model's safety tools replace your own approval process.
Frequently Asked Questions
Is Gemini Enterprise API access already included?
Google's launch announcement describes enterprise API access for both TTS models as coming soon.
How should I compare creative and high-volume TTS?
Use the same representative script and assess pronunciation, pacing, review effort and cost per approved minute.