AI Models & Platforms
ElevenLabs Launches Eleven V4 With Low-Latency Turbo Variant

ElevenLabs on September 28, 2026 launched Eleven v4, a text-to-speech model the company describes as its most emotive yet, and Eleven v4 Turbo, a low-latency variant designed for voice agents and real-time use. Both models are available immediately in ElevenAgents, ElevenCreative, and through ElevenAPI, with access included on a free account tier.
The launch post was written by Mati Staniszewski and Piotr Dabkowski and filed in the company’s Research category.
A New Architecture Focused on Expressiveness
ElevenLabs says Eleven v4 is built on an entirely new architecture designed to interpret tone, pacing, emotion, character, and context from text, generating speech that can sound dramatic, tender, urgent, comedic, or conversational while maintaining the speaker’s identity. The company says underlying audio fidelity is higher across the board than in its earlier models.
A new method for capturing speaker identities is designed to keep each voice consistent across agent conversations, audiobooks, and ads, according to the company, and scene-level context lets speakers respond to what was just said in a conversation, producing dialogue the company describes as more natural than assembling isolated lines.
Inline Audio Tags Replace SSML Controls
Users can direct delivery in natural language and embed inline audio tags; company examples include laughs, said angrily in French accent, light rain, and phone buzzing. ElevenLabs says the model follows these audio tags and direction prompts more accurately than its prior models. Support for International Phonetic Alphabet phonemes was significantly improved so custom pronunciations behave more reliably, and the ElevenLabs developer documentation covers the full tag syntax for API use. According to the product page FAQ, SSML tags such as <break> are disabled in Eleven v4, with natural-language tags serving as the control mechanism instead.
Turbo Variant and Latency Measurements
For real-time use, ElevenLabs says its research and engineering teams optimized Eleven v4 Turbo together with the ElevenAgents conversational platform as one system. Turbo supports bidirectional streaming, so audio starts returning before a sentence finishes.
The announcement’s footnoted methodology states that Eleven v4 Turbo posted a median time to first speech of roughly 150 milliseconds in September 2026 tests run over WebSocket streaming with identical scripts and default settings against Cartesia Sonic 3.6, xAI TTS, Google Gemini Flash-Lite TTS, and OpenAI GPT-4o mini TTS, with network latency measured and removed for all systems. The comparison chart on the company’s product page lists 150ms as the reference figure against a stated 262ms for Cartesia Sonic 3.6 and 814ms for OpenAI GPT-4o mini TTS. The announcement separately cites a median inference latency of about 100ms for Turbo, a figure the company says is faster than the average pause between two people talking.
Languages, Voice Cloning, and Long-Form Audio
Both models support more than 90 languages. The company says a voice recorded in one language now speaks others fluently while adopting the accent of a native speaker, and that accent adherence no longer drifts back toward the source accent over the course of a generation.
Instant Voice Clones can now capture a voice with high fidelity from 10 seconds of audio, according to the company, which also says they outperform the Professional Voice Clones of its Multilingual v2 model. Professional Voice Clones, which were unavailable in Eleven v3, are supported again and perform with the model’s full emotional range. Every clone, instant or professional, requires verified consent from the voice’s owner, and generated audio is covered by ElevenLabs’ AI Speech Classifier technology, which the company says can detect it as AI-generated.
Request stitching, the chaining of generations together for longer-form content, is significantly more reliable in Eleven v4, the company says, improving the experience of working in ElevenLabs Studio and the ElevenLabs Reader App. A single generation supports up to 10,000 characters, with context stitching keeping pacing and delivery consistent across longer scripts.
Company-Reported Benchmark Results
ElevenLabs reports that Eleven v4 ranked first on the Artificial Analysis Provider Voice Arena Leaderboard for September 2026 and was preferred by roughly 75 percent of listeners in blind head-to-head tests. The footnoted methodology says those September 2026 tests pitted the model against Cartesia Sonic 3.6, Inworld TTS-2, Google Gemini 3.8 Flash-Lite TTS, and Google Gemini 3.8 Flash TTS, with graders hearing the same line from Eleven v4 and one competitor presented blind, judging which was more expressive and which sounded more natural, and ties counted as half. A chart in the announcement states Eleven v4 won 65%–81% of those matchups.
API Access, Pricing, and Compliance
Both models are accessible through REST and streaming endpoints with TypeScript and Python SDKs, and developers switch between models with a single model_id. Output formats match every other ElevenLabs model: MP3, uncompressed WAV/PCM for studio and post-production work, and µ-law for telephony and call-center integrations, with sample rate and bitrate set through the API.
Pricing follows the same credit structure as the company’s other text-to-speech models: a free tier with 10,000 credits per month, which the company equates to roughly 10 minutes of audio; paid plans starting at $6 per month that include professional voice cloning; and custom enterprise pricing. Every voice in the company’s library of more than 17,500 voices works with Eleven v4, though instant and professional clones created before the launch need to be retrained to work effectively with the new model.
ElevenLabs states that Eleven v4 runs on infrastructure certified SOC 2 Type II, ISO 27001, and PCI DSS Level 1, is GDPR compliant with HIPAA-eligible workflows for healthcare, does not train on scripts or audio tags without consent, and offers Zero Retention Mode for eligible enterprise services.
The Eleven v4 product page carries statements from named customers, including Ryan Peterson, SVP of product for Agentforce Voice at Salesforce, who said: “That’s exactly where we’re seeing Eleven v4 Turbo raise the bar, with faster, more natural responses that meet the standard our customers expect.” BeyondWords co-founder Patrick O’Flaherty, Spring Financial head of credit building products Oscar Daniels, and Accenture Accelerate music and audio lead Kyle Gudmundson also provided statements on the new models.
ElevenLabs directs developers to its documentation for the full audio-tag syntax and API integration details.












