· via The Verge
Google's Gemini 3.5 Transcribe removes filler words and handles jargon in 85+ languages
Google's new Gemini Audio models strip filler words from transcripts, recognize specialized jargon and 85+ languages, and are rolling out on macOS, Android and in developer preview.

Google has expanded its Gemini Audio lineup with three new models — Gemini 3.5 Transcribe, 3.5 Live, and 3.5 Live Experimental — that promise noticeably smarter speech handling, including transcripts that automatically drop filler words such as "um" and "uh." According to The Verge, the rollout began on August 26.
What's new
The centerpiece is Gemini 3.5 Transcribe, an entirely new model rather than an incremental upgrade to Google's existing speech stack. The Verge reports that Google positions it as a major step up from its previous transcription model, Chirp 3, with particular gains in multilingual performance and word error rates. It supports more than 85 languages and is built to cope with the messy realities of real speech: background noise, interruptions, and specialized vocabulary.
Cleaner transcripts by default
The standout capability is automatic cleanup. Beyond converting speech to text, 3.5 Transcribe can format the output and remove filler words, so a transcript reads more like edited prose than raw dictation. Google also lets users supply a custom vocabulary, meaning names, product terms, or domain-specific jargon can be recognized and spelled correctly without manual correction — a long-standing pain point with traditional dictation tools.
Users can also "edit naturally with just your voice," per Google, revising the transcript without touching a keyboard. For pre-recorded audio, the model can attribute speech to up to three distinct speakers and produces word-level timestamps, which makes the output easier to navigate, search, and repurpose.
The Live models
Alongside Transcribe, Google shipped Gemini 3.5 Live and Gemini 3.5 Live Experimental, both building on the speech recognition technology behind Gemini's voice chat mode. According to The Verge, 3.5 Live improves on handling mid-sentence interruptions, language recognition, and live visual processing. The Experimental variant goes further: it narrates its own progress step by step in real time while working through more complex reasoning tasks, giving users visibility into what the model is doing rather than leaving them waiting on a silent final answer.
Where you can get it
The update is reaching all macOS Gemini app users in English, plus the Rambler dictation feature on Android in select countries and languages. Developers can try the models in public preview through the Gemini API via AI Studio and Antigravity. Chrome support is listed as coming soon.
Notably, The Verge points out that these audio models arrive while Google's Gemini 3.5 Pro — which the company said would launch in June — has yet to materialize.
Why it matters
Transcription is among the most widely used AI features, and its weaknesses are well known: mangled names, missing punctuation, verbatim filler words, and poor handling of accents or technical jargon. Automatic filler removal shifts AI transcription from a raw record toward something closer to a first-draft document, which could meaningfully cut cleanup time for journalists, students, and anyone working from meeting notes.
Custom vocabulary support and multi-speaker attribution target the remaining gaps that make many transcripts unreliable in practice. And with the models available in public preview, developers can embed these capabilities into their own applications without waiting. The continued absence of Gemini 3.5 Pro is worth watching, but on the audio side Google is clearly pushing dictation and voice interfaces closer to everyday, dependable use.
- #gemini
- #transcription
- #speech-recognition
- #voice-ai