The part I care about isn’t the fire-breathing dragon demo. It’s the stage directions. You write a script, mark where the voice sighs or says mhm, and the model performs it. That turns TTS from a settings panel into something closer to a prompt, which is how I already work with Claude all day.
What I’d actually do: render a Cloudy Brain post with Flash TTS and check whether the long-form claims hold over a full read without the narrator drifting. Voice cloning is the tempting bit, but it needs a verbal consent recording from the voice owner first, which is the right amount of friction. Stock voices it is, all 2,000-plus of them.
The story — Google introduced Gemini 3.8 Flash TTS, built for voice design and line-by-line performance direction, and Flash-Lite TTS, aimed at high-volume dubbing and voice agents. Flash TTS creates voices from natural language prompts across more than 100 languages, offers 2,000+ library voices, and replicates a voice from a 30-second sample with consent verification. All output carries SynthID watermarks. Both are rolling out in the Gemini API and Google AI Studio. (Source)