· via dev.to (home feed)
Microsoft's MAI-Transcribe-2 covers 60 languages with claimed best accuracy at $0.10 per hour
Microsoft AI's MAI-Transcribe-2 transcribes 60 languages, tops the FLEURS benchmark in its own tests, and undercuts OpenAI, Google and ElevenLabs on price at $0.10 per hour of audio.

On September 3, 2026, Microsoft AI released MAI-Transcribe-2, its latest speech-to-text model, covering 60 languages and claiming the top spot on a widely used benchmark. According to a dev.to summary of the announcement, the company positions the model as the fastest, most accurate and cheapest transcription option on the market, at an introductory price of $0.10 per hour of audio.
What the benchmarks show
Microsoft's headline claim rests on FLEURS, a benchmark published by Google researchers in 2022 in which native speakers read roughly 2,000 sentences across 102 languages. MAI-Transcribe-2 reportedly achieves an average word error rate of 5.2% across its 60 supported languages, which Microsoft says ranks first.
For Thai specifically, the model posts a 3.4% word error rate, the best figure in Microsoft's comparison table, ahead of Gemini 3.5 Transcribe at 3.8%, Gemini 3.1 Pro at 4.7%, OpenAI's GPT-transcribe at 5.4%, ElevenLabs' Scribe v2 at 6.1% and Whisper v3-large at 8.7%.
Independent measurements add nuance. Artificial Analysis, which tests models through public APIs, places MAI-Transcribe-2 second on its word error rate leaderboard while noting that it sits on the Pareto frontier for accuracy versus latency, meaning no competitor is more accurate without being slower, or faster without being less accurate. On raw speed, Microsoft says the model is roughly 10x faster than OpenAI's GPT-Transcribe, 7x faster than ElevenLabs' Scribe v2 and 5x faster than Google's Gemini 3.5 Transcribe, based on Artificial Analysis evaluations.
New capabilities
The language count has grown from 25 in the original April release and 43 in the June v1.5 update to 60 now, making this Microsoft AI's third transcription model in five months. The new version also adds features aimed at production workloads:
- Speaker diarization, which attributes text to individual speakers in multi-person recordings
- Word-level timestamps, useful for search, editing and syncing subtitles to video
- Keyword biasing, letting developers supply lists such as drug names, product codes or employee names so specialised terms are transcribed correctly
- Code-switching support for conversations that mix languages mid-sentence, such as Hinglish and Spanglish
- Two output modes: a verbatim mode that keeps filler words and hesitations for legal and compliance work, and a clean mode that strips them for readable captions
Pricing and the cost argument
The $0.10 per hour launch price is promotional through the end of the year and represents a drop of roughly 72% from the $0.36 per hour charged for the first model five months earlier. VentureBeat laid out the enterprise implication: an organization transcribing 100,000 hours of call center audio per year, a volume typical for a large bank or telecom, would see its bill fall from $36,000 to $10,000, turning transcription into a routine line item rather than a budget debate.
As the dev.to write-up notes, the high throughput also underpins the pricing itself, since a faster model consumes fewer GPU hours for the same workload. Developers can try the model through Microsoft Foundry, MAI Playground and OpenRouter.
Strategic backdrop
The release fits a broader Microsoft pattern of building its own frontier models one modality at a time and gradually swapping them in for OpenAI technology across its products, with transcription reportedly moving fastest. Microsoft AI CEO Mustafa Suleyman told The Verge that the first transcription model came from a roughly ten-person team kept free of bureaucracy, running at about half the GPU cost of comparable frontier models.
Caveats worth noting
The benchmark figures are vendor-reported, and the strongest numbers come from Microsoft's own comparison tables. FLEURS in particular uses read speech from native speakers rather than real-world audio with background noise, crosstalk and overlapping speakers, so production accuracy will depend heavily on the deployment scenario.
Why it matters
If the claims hold up, transcription at $0.10 per hour with benchmark-leading accuracy effectively turns audio-to-text into near-free infrastructure for meetings, subtitles, media archives and call center analytics, and the strong Thai result shows the gains reaching well beyond English. For OpenAI, Google and ElevenLabs, it means Microsoft is now competing directly against its own partners on both price and speed. It is also another marker in Microsoft's shift away from relying on OpenAI toward its own in-house model stack, with speech recognition emerging as the first front where that transition is visibly ahead.
- #microsoft
- #speech-to-text
- #ai-models
- #transcription
- #benchmarks