· via Google AI blog
Google says its AI now handles 300+ languages spoken by 86% of the world's population
Google says its products and AI models now operate in more than 300 languages spoken by over 7 billion people. The company detailed new speech models, on-device translation and data partnerships aimed at underrepresented languages.

Google says its AI technologies now power everyday interactions in more than 300 languages, spoken by over 7 billion people and representing 86% of the global population. The company laid out the milestone, and the work behind it, in a September 15 post on the Google AI blog, arguing that technology has long worked best for a handful of dominant languages while leaving thousands of others poorly served online. All of the figures below come from Google's own announcement.
Speech models that skip the transcript
According to Google, conventional speech systems have relied on a rigid pipeline: transcribe audio to text, process the text, then synthesize it back into sound. That design strips out tone, pacing, emotion and context, and struggles with how people actually talk — hesitations, overlapping speech, and mixing languages such as Spanglish or Hinglish within a single sentence. To capture this, Google says it has shifted to training models like Gemini to work with audio directly, grasping both sound and intent.
The blog names Gemini 3.5 Live Translate, which provides real-time spoken translation across 70 languages and more than 2,000 language pairs while handling code-switching and emotional cues. Its counterpart, Gemini 3.5 Transcribe, is described as Google's most accurate speech-to-text model yet, producing polished and formatted text even in noisy environments or with specialized jargon. Transcribe also powers a Gboard feature called Rambler on Android, which removes filler words, fixes grammar and punctuation, and lets users edit or rewrite text by voice while switching between languages.
Under its 1,000 Languages Initiative, Google aims to support the world's 1,000 most-spoken languages. Its Universal Speech Model was trained on 12 million hours of audio and uses cross-lingual transfer learning, which applies patterns learned from data-rich languages to those with far less training data. Google says the effort rests on 25 years of research and more than 400 peer-reviewed papers on speech.
Collecting data through local partnerships
Because the web disproportionately reflects a few dominant languages, Google says it had to rethink how it gathers data, turning to grassroots partnerships. Its open-data projects include WAXAL (Wolof for "speaking"), built with Makerere University and Digital Umuganda, a speech dataset spanning 27 Sub-Saharan African languages spoken by over 100 million people in more than 26 countries. Project Vaani, run with the Indian Institute of Science and Bhashini, organizes collection by region rather than by language, and has so far gathered over 30,000 hours of speech in 109 languages from more than 155,000 speakers.
A third effort, the Amplify Initiative, enlisted more than 1,600 local experts and 20 universities across four continents to contribute 15,000 multimodal data points capturing local nuance. Google also introduced Language Explorer, a visualization tool for LinguaMeta, which it describes as the largest open-source language data repository in the world, continuously mapping more than 7,000 spoken, written and signed languages. Google.org separately backs the Centre for Digital Language Inclusion and AI Singapore's Project Aquarium, efforts to bring multilingual tools to farmers, healthcare workers and teachers.
Translation without connectivity
Google notes that reliable internet remains out of reach for more than 3 billion people. Its software answer is TranslateGemma, a family of lightweight open translation models derived from Gemini and trained on 55 languages, which run on-device so that translation no longer needs a cloud connection. For the hundreds of millions still on feature phones, Google is supporting Viamo's voice assistant AVA, which brings Gemini to basic handsets; after a pilot in Rwanda with interactive voice response users, the service has fielded more than 2 million Gemini-powered questions.
Sign language and local pronunciation
On accessibility, Google unveiled Sign Language-to-Text (SL2T), trained across more than 50 sign languages. It enables sign-to-text dictation in Gboard and Live Transcribe on the Pixel 11, beginning with American Sign Language to English — an early step toward serving the roughly 70 million people worldwide who rely on sign language. In New Zealand, Google worked with Māori language experts to improve how Maps pronounces place names, folding culturally authentic pronunciations into its text-to-speech models.
Why it matters
Most language technology has been built for a small set of dominant languages, leaving billions of speakers with tools that partly or entirely miss how they communicate. If Google's coverage claims hold up beyond the company's own accounting, the combination of community-collected data, on-device models and feature-phone access would extend AI to populations that cloud-first tools have largely skipped. The announcement also signals where competition in AI is heading: raw model quality is becoming table stakes, and reach — across languages, devices and connectivity levels — is turning into the differentiator. As with any vendor milestone post, the numbers are unaudited, but the shift toward multilingual, offline-capable AI is a trend worth tracking.
- #translation
- #speech-recognition
- #multilingual-ai
- #accessibility