· via Hacker News – Front Page (native)
Google launches Gemini 3.8 Live voice models that reason while speaking
Google has released Gemini 3.8 Live and 3.8 Live Extended Thinking, speech-to-speech models that reason while talking, lead several voice AI benchmarks and roll out across Search, Workspace and the Gemini API.

Google has introduced two new live dialogue models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, according to a September 15 announcement on the company's blog that reached Hacker News's front page. The models are built for near real-time reasoning and aimed at two audiences at once: developers assembling production voice agents, and consumers using voice features in the Gemini app, Google Workspace and Search.
Two models, two jobs
According to Google, the standard Gemini 3.8 Live is designed for scale and cost efficiency, pairing conversational intelligence with fluid dialogue and visual grounding. The Extended Thinking variant targets high-complexity work that needs multi-step reasoning, which Google describes as delivering enterprise-grade task completion and intelligence.
The signature capability is reasoning and speaking at the same time. Rather than going silent while it computes an answer, Extended Thinking acknowledges a prompt with early verbal cues such as "Let me check that…" and narrates its progress as multi-step tasks run in the background, keeping the conversational flow uninterrupted.
The standard model processes visual input in near real time to give responses more context, automatically detects and switches between 97 supported languages mid-conversation, and can execute tools and API calls in the background while the user keeps talking.
Benchmark claims
Google reports that Gemini 3.8 Live Extended Thinking took first place overall on Artificial Analysis' Speech to Speech Quality Index with a score of 82.6. The company says it also leads on agentic task completion, scoring 68.6% on τ-Voice and 35.1% on Sierra's τ-Voice-banking benchmark, and posts 97.7% on Big Bench Audio for reasoning while remaining price-competitive with other frontier models.
The base 3.8 Live model placed second in the Speech Agent Arena, which Google cites as an indication of user preference. On ServiceNow's EVA-Bench, a benchmark for evaluating voice agents, Google says the models push the Pareto frontier for complex workflows by balancing accuracy with conversational quality. The company notes those runs were conducted on the Live API on the Gemini Enterprise Agent Platform.
Where it is rolling out
Both models started rolling out on announcement day. Developers get them in the Gemini API and Google AI Studio. Enterprises get a private preview in Gemini Enterprise, with availability in Gemini Enterprise for Customer Experience described as coming soon.
For consumers, 3.8 Live is available in Search Live, where Google says it can walk users through step-by-step, real-time troubleshooting. Extended Thinking is reaching Gemini Live, Google AI Pro and Ultra subscribers in Workspace in Docs, and all Google AI subscribers in Gmail and Keep.
Ecosystem and safety
Google says developer platforms including Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel and Vision Agents are building on the Gemini Live API, handling real-time media streaming infrastructure so developers can focus on the user experience. Salesforce, Genspark and Lumeris are named as early partners, with Google highlighting the models' latency, fluidity and tool-calling capabilities.
The company also states that all audio generated by its AI products carries a SynthID watermark woven into the output, intended to keep synthetic speech detectable and help prevent misinformation. The announcement page itself carries a note stating that the content is generated by Google AI.
Why it matters
The gap between typing to a chatbot and talking to one has mostly come down to latency and reasoning depth: voice assistants have historically either paused awkwardly while thinking or gave shallow answers to hard questions. A model that reasons while it speaks, narrates background work and fires off tool calls mid-conversation makes voice a viable interface for genuinely complex tasks rather than just quick queries. Competitive pricing on top of that lowers the barrier for teams building production voice agents. If the benchmark claims hold up under independent scrutiny, Gemini 3.8 Live sets the reference point for what a speech-to-speech model should deliver, and it puts pressure on rivals in a market where real-time voice is quickly becoming the next battleground.
- #gemini
- #voice-ai
- #speech-to-speech
- #ai-models