
What happened
ElevenLabs launched Eleven v4 and low-latency Eleven v4 Turbo, ranked #1 by Artificial Analysis and preferred by ~75% of listeners in blind tests over competitors like Cartesia Sonic 3.6 and Google Gemini models.
Why it matters
ElevenLabs says the new architecture produces voiced dialogue that sounds emotionally directed rather than mechanically read aloud, which it positions as a shift from earlier text-to-speech generations.
What to watch
The ranking rests on blind preference tests run in September 2026, so the test is whether third-party evaluations repeat those results now that both models are shipping through ElevenAgents, ElevenCreative, and the API.
WHO IT HITSProduct and localization teams who currently ship voice agents, audiobooks, or dubbed content face a new option: ElevenLabs is pitching Eleven v4 as emotive enough for character work and Eleven v4 Turbo as fast enough for live agents, though switching costs and integration testing will determine adoption.
Summaries like this, in your inbox every morning.
ElevenLabs frames Eleven v4 and Eleven v4 Turbo as the culmination of its latest research in expressive speech generation, built for content where delivery matters as much as the words. The company positions the models against a specific trade-off it says has shaped the market: high-quality voice models have tended to be slower, pushing many builders toward proficient but monotone agents. Eleven v4 Turbo is aimed squarely at that gap, pairing speed with emotion, and it is optimized together with ElevenLabs' ElevenAgents platform as one system rather than stitched from separate vendors.
Beyond delivery, the launch addresses consistency and multilingual reach. Eleven v4 uses a new method for capturing speakers' identities, preserves voice consistency across generations and long-form projects like audiobooks and ads, and lets a voice recorded in one language speak any other fluently while adopting a native accent. Instant Voice Clones now need just 10 seconds of audio, and Professional Voice Clones are supported for high-fidelity needs. Inline tags and natural-language direction prompts, from [laughs] to [said angrily in French accent], are followed more accurately than in prior models.
The stakes hinge on whether the blind-test advantage holds once customers can run the models themselves. The #1 ranking and ~75% listener preference come from September 2026 tests against named competitors including Cartesia Sonic 3.6, Inworld TTS-2, and Google Gemini models, so independent replication and real-world deployments in agents, dubbing, and audiobooks will determine whether ElevenLabs' emotive claim translates into lasting adoption.
For example, today's edition would include:
AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Modulate raised $25 million, led by Future Ventures with returning investors Hyperplane and Lakestar, bringing…
Tarini Padmanabhuni founded DetectifAI after a deepfake of her grandfather's brother's voice tricked him into…

Google said its Gemini 3.8 Live voice model now has a Live Avatar feature, and that Live Avatar-equipped Gemin…

Google announced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, text-to-speech models that recreate a voi…

OpenAI said on 9月24日 that ChatGPT Voice on mobile and web can now handle work using connected apps and tools…

Notta's "Notta Memo Pro" is an AI voice recorder (a device that records and automatically converts speech to t…
