AI RACE— The AI Race
New Models

ElevenLabs Launches Eleven v4 Speech Model and 150ms Turbo Variant

ElevenLabs has released its Eleven v4 speech model alongside an ultra-low-latency Turbo edition, delivering enhanced emotional direction, expanded multilingual cloning across 90 languages, and 150-millisecond response times.

09/29/2026, 21:45
New Models

ElevenLabs Debuts Eleven v4 and Real-Time Turbo Model

Voice AI developer ElevenLabs has officially released Eleven v4, a new flagship text-to-speech model engineered to follow directional cues more faithfully and maintain voice consistency over long audio productions. Alongside the standard release, the company unveiled Eleven v4 Turbo, a lower-latency variant tailored for interactive voice agents, video game characters, and customer service workflows.

The update builds on Eleven v3, which debuted roughly a year prior. While v3 introduced scripting tags for audio behaviors like whispers, laughter, and sound effects, ElevenLabs says v4 executes these tags far more reliably while improving character consistency across regenerated lines. Both models are immediately accessible across ElevenCreative, ElevenAgents, and the company's API.

Architecture Enhancements, Dubbing Controls, and Benchmark Gains

Eleven v4 features a revised architecture designed to interpret the broader context, tone, and pacing of a full scene rather than synthesizing sentences in isolation. Users can provide directions via inline tags or natural-language sentences, with upgraded phonetic spelling controls to dictate the exact pronunciation of names and technical terms. In internal pronunciation evaluations, v4 scored 91.7 percent, up from the 85.6 percent mark achieved by v3.

The model processes up to 10,000 characters per single request—roughly ten minutes of spoken audio—allowing audiobooks and extended productions to maintain consistent delivery across segment boundaries. In multi-speaker dialogue, characters now adapt dynamically to conversational context without timbre drift.

Language coverage has expanded from approximately 70 languages in v3 to more than 90 in v4. The system allows cloned voices to speak foreign languages with native accents while avoiding accent deterioration over longer takes. ElevenLabs has also reinstated its Professional Voice Clones feature, which was unavailable in v3, alongside its Instant Voice Clone tool, which requires ten seconds of sample audio. Actors licensing their voices through ElevenLabs' marketplace—which already includes talent like Michael Caine—can monetize extensively trained replicas for global dubbing across all supported languages.

On the Artificial Analysis Provider Voice Arena leaderboard, Eleven v4 currently ranks ahead of Cartesia Sonic 3.6 and Google's Gemini 3.8 Flash TTS. In blind evaluations conducted by ElevenLabs, roughly 75 percent of listeners preferred v4 over models from Cartesia, Google, and Inworld, with v4 rated as more expressive in 65 to 81 percent of head-to-head comparisons depending on the competing system.

Turbo Latency, Pricing Structure, and Regional Processing

To address the trade-off between latency and expressive output in real-time applications, Eleven v4 Turbo was co-optimized alongside the company's ElevenAgents platform. In ElevenLabs' testing, v4 Turbo delivers audible speech in 150 milliseconds. By comparison, Cartesia Sonic 3.6 clocked in at 262 milliseconds, while OpenAI’s GPT-4o mini TTS registered 814 milliseconds.

Standard API pricing is set at $80 per million characters for Eleven v4 and $40 per million characters for v4 Turbo. ElevenLabs is running promotional pricing through October 12, lowering the rates to $22 and $11 per million characters, respectively. Subscribers on the $22-per-month Creator plan or higher can test v4 in ElevenCreative at no additional charge for two weeks, subject to a cap of twice their monthly credit allotment. For market comparison, Artificial Analysis lists Cartesia Sonic 3.6 at $49 per million characters and Gemini 3.8 Flash TTS at $16.49 per million characters.

Customer audio data is stored in the United States by default. Enterprise clients have access to isolated data storage in the European Union, India, or Singapore, though some operational processing may still occur externally; EU customers can enforce domestic API processing by enabling a zero-data-retention mode. The v4 release follows ElevenLabs' mid-September rollout of Music 2.5, a model dedicated to generating denser, more natural musical tracks.

◗ Sources

The Decoder09/29

Related stories