Google DeepMind has introduced two new text-to-speech models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, positioning them as a shift from static voice presets to a more dynamic creative tool. The Flash model is aimed at deep creative direction and character design, while Flash-Lite targets high-volume, cost-efficient use cases like dubbing and expressive voice agents.

Both models support creating entirely new voices from natural language prompts, with control over role, accent, and vocal characteristics across more than 100 languages and dialects. They also allow line-by-line direction of performances, including pacing, emotion, dialect shifts, and conversational sounds such as laughs or interjections. The models can stage two-speaker scenes from a single script and maintain voice consistency over hours of audio, which the company says suits audiobooks and podcasts.

For voice replication, a 30-second audio sample can be used to recreate a consistent vocal profile, backed by consent verification, SynthID watermarking, and C2PA credentials. The models are available across Google AI Studio, the Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids. Since this is a single announcement, there are no independent sources to compare, but the company frames these features as expanding the Gemini Audio family following earlier releases like 3.5 Live Translate and 3.8 Live Extended Thinking.