· via TechCrunch
ElevenLabs v4 and v4 Turbo add stackable expression tags, 90+ languages and lower latency
ElevenLabs has released v4 and v4 Turbo, its latest speech models, bringing 10-second voice cloning, expandable expression tags and support for over 90 languages, with latency cuts aimed at voice agents.

What launched
ElevenLabs introduced two new speech models on Monday, ElevenLabs v4 and v4 Turbo, TechCrunch reports. The release follows the company's v3 model from last year, which it teased at an event in Warsaw earlier this year. The headline improvements are more granular control over vocal expression, reduced latency aimed at conversational agents, and a jump from around 70 supported languages to more than 90.
The v4 generation runs on a redesigned architecture that the company says delivers both finer control and quicker voice cloning. According to TechCrunch, users can now clone a voice from just 10 seconds of audio.
Expression control and longer-form speech
On the creative side, the new model is built to hold a consistent voice identity across longer stretches of text, a common weak point for speech synthesis. It also reads with an awareness of surrounding context, shifting its delivery to match what the text actually says rather than applying a uniform tone.
The biggest workflow change is in how expression is directed. ElevenLabs introduced inline tags for steering delivery with v3, and v4 broadens that system: users can now stack multiple tags together, and the model respects the sequence in which they appear. That should let creators script more layered performances, such as moving from whispered to urgent within a single passage, without regenerating audio repeatedly.
On languages, the company told TechCrunch the most noticeable quality gains came in Japanese, Brazilian Portuguese, Mandarin and Cantonese.
Built for voice agents
A large part of the pitch targets enterprise calling, a business ElevenLabs says has scaled quickly over the past year, with more than 55% of its revenue now coming from large companies. For that use case, v4 lowers latency so conversations feel more natural, and it can begin producing audio as soon as the underlying language model starts emitting its answer, instead of waiting for the full response to finish generating.
TechCrunch also reports that the model handles tense moments differently: it is tuned to deal with confrontations, escalations and hold requests in ways intended to improve how issues get resolved on calls, a nod to the reality that agent conversations do not always go smoothly.
A crowded and well-funded field
The launch lands amid intensifying competition. Startups including Cartesia, Deepgram, Fish Audio, Boson and WellSaid Labs have all built expressive speech models, while larger players such as Google and OpenAI have kept improving their own voice offerings.
ElevenLabs' commercial position appears strong regardless. The company raised $500 million in a Sequoia-led round earlier this year at an $11 billion valuation, and TechCrunch reports rumors of a subsequent round that would value it at $22 billion. Its annualized revenue run rate has reportedly climbed from roughly $330 million at the start of the year to more than $600 million, and headcount has passed 800, with hiring across India, Europe and Brazil. In a recent interview with TechCrunch, co-founder and CEO Mati Staniszewski said the company is aiming for an IPO "in the next years," without committing to a timeline.
Why it matters
Voice is quickly becoming the primary interface for AI agents, and the competition is no longer just about intelligibility. The differentiators are now latency, emotional range and reliability under real conversational pressure, which is exactly where ElevenLabs is aiming v4: streaming audio in parallel with the LLM's output, stacking expressive directions, and handling difficult customer moments all point at live agent calls as the main battleground.
The language expansion matters for the same reason. Enterprise deployments are global, and significant gains in Japanese, Mandarin and Portuguese make the model viable in markets where earlier generations lagged. Combined with the company's reported revenue growth and IPO ambitions, this release is less about a demo-quality party trick and more about capturing the infrastructure layer for voice-based AI, ahead of both well-funded startups and the largest AI labs.
- #ai
- #text-to-speech
- #voice-agents
- #elevenlabs
- #speech-synthesis