· via Hacker News – Front Page (native)
Google launches Gemini 3.8 text-to-speech with voice cloning and line-by-line control
Google has released Gemini 3.8 Flash and Flash-Lite TTS models, adding prompt-based voice creation, consent-gated cloning and line-by-line performance control across its AI platforms.

Google has launched two new text-to-speech models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, presenting them as a significant step up in expressiveness for its speech stack. According to Google's announcement, which surfaced on the Hacker News front page, the models are available across Google AI Studio, the Gemini API, Gemini Enterprise, Gemini Notebook and Google Vids.
Two tiers: creation and scale
Google positions Flash TTS as the creative tier, built for character design and detailed direction. It can generate original voices from natural-language descriptions — specifying a role, an accent and other vocal characteristics — across more than 100 languages and dialects, and it allows line-by-line control over acting cues, pacing, dialect shifts and backchanneling.
Flash-Lite TTS is tuned for high-volume, cost-efficient workloads such as dubbing, audio content production and expressive voice agents, while keeping fine-grained control over tone and pacing, the company says.
Both join an audio lineup that already includes Gemini 3.5 Live Translate, 3.5 Transcribe, 3.8 Live and 3.8 Live Extended Thinking.
From 30 presets to a vocal studio
Google says users can scale up from the 30 voices the platform previously offered to an effectively open-ended library. Alongside generative voice design, users get access to more than 2,000 production-ready voices covering regional varieties such as Mexican Spanish, Quebec French and Scots English.
Voice replication is also included: a consistent vocal profile can be rebuilt from a 30-second audio sample. Google says the feature is gated by consent verification, requiring a verbal consent recording from the voice owner that matches the reference speaker, and that outputs carry SynthID watermarking and C2PA credentials. A voice remixing option for adjusting timbre, pitch, pace and accent on an existing voice is listed as coming soon.
Direction, dialogue and non-verbal cues
Both models support long-form generation, which Google claims can hold voice quality and character timbre across hours of continuous audio with minimal speaker drift — a pitch aimed at podcasts and audiobooks. A native two-speaker mode stages multi-turn conversations from a single script while keeping the two voices distinctly separated, and users can insert scripted non-verbal cues such as laughs, sighs and gasps, along with active-listening interjections, to shape reaction beats and timing.
Benchmark results
Google reports that Flash TTS took first place on Hume AI's Voice Design Benchmark with a score of 71.4, and also led that benchmark's accent modelling category at 60.8. The two models placed first and second on Hume AI's Overall Quality Index. In blind human-preference tests on Voice Arena, Google says they ranked at the top in languages including Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish and Hindi, with improvements over the earlier Gemini 3.1 Flash TTS particularly in long-form content and dual-speaker scripts. These figures come from Google's own announcement rather than independent evaluation.
Where to try it
A new audio playground in Google AI Studio, available now, combines voice design and replication with a dual-speaker screenplay editor for directing delivery line by line. Through the Gemini API, developer platforms including Agora, LiveKit, Pipecat and Vercel are wiring the models into their stacks, and Google lists Figma, HeyGen, Linguana, Wondercraft, 99.co and Ollang among partners applying them to dubbing, localisation and conversational voice.
Why it matters
Speech is rapidly becoming the default interface for AI products, and Google is pushing text-to-speech beyond fixed presets toward prompt-driven performance. Features such as per-line direction, multi-speaker staging and stable long-form output target production workflows — audiobooks, dubbing, games, agents — where current tools often fall short. The consent checks, SynthID watermarking and C2PA credentials are an attempt to get ahead of the voice-cloning and misinformation concerns the technology attracts. For developers and studios, the open question is how these models compare on price and real-world quality against established commercial speech providers; Google's benchmark numbers are self-reported, so hands-on testing in AI Studio will be the real measure.
- #gemini
- #text-to-speech
- #ai-audio
- #voice-cloning