Google introduced Gemini 3.8 Flash TTS and Flash-Lite TTS on September 23, 2026. The releases add custom voice design, line-by-line direction, and multi-speaker speech generation. They concern text-to-speech output, a different component from speech recognition or a complete conversational agent.
The announcement describes developer rollout through the Gemini API and AI Studio, while enterprise API distribution is described as coming soon. The audio model card supplies limitations and safety context.
Custom voices need a controlled production process
Google describes designing voices and directing delivery for individual lines. That may be useful for narration, localization, and a consistent product voice. A team should define pronunciation, tone, pacing, and the intended audience before selecting a voice.
Evaluate a complete script, not only a short demonstration. Names, numbers, abbreviations, and domain terminology can reveal problems that an expressive sentence does not. Retain the configuration and approved output so later revisions can be compared consistently.
Dialogue generation is different from a live conversation
The launch includes native two-speaker staging and longer-form generation. Producing a convincing dialogue track does not mean the model has completed an external task or verified every statement in the script.
Keep content approval separate from audio approval. Check the text first, then examine speaker consistency, pauses, and pronunciation. If a user interacts with generated speech, the application also needs to decide when speech can be interrupted and how a corrected answer replaces the original.
Consent and provenance belong in the workflow
Google says voice replication requires a matching verbal consent recording from the voice owner. It also says Gemini Audio output carries SynthID watermarking. These are specific product controls, not a general permission to recreate anyone’s voice.
A business should retain the relevant authorization and know the scope of permitted use. A voice approved for one campaign may not be approved for every later context. Watermarking also does not establish the factual accuracy of the spoken content or resolve every disclosure requirement.
What to test before deployment
The model card acknowledges general foundation-model limitations and possible slowness or timeouts. Measure output quality, delivery time, failures, and recovery. If speech is generated dynamically, ensure a stalled request does not leave the user without a clear status.
Nerova’s assessment is that the release gives creators and developers more control over vocal performance. The best deployment combines approved scripts, explicit voice rights, reproducible settings, and realistic reliability checks. A strong voice interface remains a complete product responsibility.