Technology

ElevenLabs v4 Brings Faster Voice Cloning, Sharper Direction, and 90-Plus Languages

Martin HollowayPublished 6d ago3 min readBased on 7 sources
Reading level
ElevenLabs v4 Brings Faster Voice Cloning, Sharper Direction, and 90-Plus Languages
source:elevenlabs.io

ElevenLabs has launched ElevenLabs v4 and v4 Turbo, two speech synthesis models with expanded expression control and support for more than 90 languages. TechCrunch

The release uses a new architecture built for better control and faster cloning. V4 can clone a voice from 10 seconds of reference audio. That short sample helps in production work, where clean recordings are scarce or must come from users.

Language coverage is the other main change. The prior v3 generation supported 70 languages. V4 moves beyond 90. ElevenLabs said the biggest quality jump was in Japanese, Brazilian Portuguese, Mandarin and Cantonese, which points to focused work on prosody — rhythm and intonation — tokenization, how text is split for the model, and data balance for those locales rather than even gains everywhere.

For control, v4 expands inline expression tags. Users can stack multiple tags and the model follows the sequence through a passage. In practice, direction can sit inside the text stream, for example shifting delivery mid-paragraph, without splitting jobs or joining separate takes later. Company documentation describes Eleven v4 as its "most emotionally rich, expressive speech synthesis model." ElevenLabs Docs

Latency gets attention for conversational use. V4 has lower latency for voice agents to allow more fluid conversation. It can start generating audio as soon as the underlying LLM starts generating answers, rather than waiting for a full response. That fits text-to-speech to LLM outputs that arrive in pieces and cuts turn-taking gaps in agent pipelines.

Documents place v4 inside a wider speech stack. Eleven v4 has a 10,000 character limit per request. Alongside synthesis, Scribe v2 is listed as the state-of-the-art speech recognition model for transcription across 90-plus languages, with precise word-level timestamps. ElevenLabs Models Other material claims Scribe outperforms Whisper, Deepgram and Gemini in benchmark tests, and cites a library of 5,000-plus voices in more than 70 languages. Those figures come from undated documentation and should be read as background to the dated v4 launch details above.

On the commercial side, ElevenLabs and Universal Music Group announced a multi-year strategic agreement on September 10, 2026. The agreement will enable fans to create remixes and mashups. ElevenLabs

The broader context here is clear for teams shipping voice systems. Controllable, low-latency text-to-speech has become the binding constraint for voice agents, dubbing pipelines and localized content work. Raw intelligibility stopped being the differentiator some time ago. What slows deployment now is directing performance at scale, keeping speaker consistency from very short prompts, and keeping first-audio delay low enough for interruption to feel natural.

Looking at what this means for implementation, three details deserve attention. First, 10-second cloning eases sign-up and misuse in equal measure, lowering friction for personalization while raising the bar for consent. Second, stacked tags allow one model run to follow instructions in order, which is useful but needs testing for accuracy, spillover, and results outside the four cited languages. Third, starting from partial output trades sentence planning for speed, helping live dialogue more than long narration.

In my view, the direction is encouraging. Finer direction, shorter enrollment and closer links to chatbot output point to speech becoming a programmable layer rather than a post-production step. Measurement, guardrails and checks per language come next, but builders have more to work with for assistants, translated media and accessible content.