September 28, 2026 — @ElevenLabs introduced Eleven v4 and Eleven v4 Turbo on X (2.2M+ views), calling them the company’s fastest and most emotive voice models yet and citing a #1 rank from Artificial Analysis on launch day. Product pages and samples live at elevenlabs.io/v4 with immediate availability in ElevenCreative, ElevenAgents, and ElevenAPI.
Launch announcement on X
Inline performance tags (demo thread)
ElevenLabs’ follow-up clip shows directing delivery in-script — emotion, pacing, SFX — without a separate DAW pass:
TL;DR — v4 vs v4 Turbo
| Eleven v4 | Eleven v4 Turbo | |
|---|---|---|
| Best for | Audiobooks, ads, character, long-form | Voice agents, support, realtime calls |
| Latency | Quality-first | ~100ms median inference; ~150ms time to first speech (vendor) |
| Architecture | New expressive stack; context stitching for long scripts | Same expressive range, optimized for stream-in/stream-out |
| Languages | 90+ (incl. Cantonese, Mongolian expansions in marketing) | Same |
| Voice clones | IVC from 10s audio; PVC restored vs v3 gap | PVC consistent across turns |
API modelId | eleven_v4 | Turbo id on docs / dashboard (switch by model_id) |
| Promo (2 weeks) | $22/M chars API; Creator+ 2× credits on v4 | $11/M chars API |
ElevenLabs’ marketing compares Turbo time-to-first-speech against Cartesia Sonic 3.6 and OpenAI GPT-4o mini TTS — always re-benchmark on your text length, voice, and region; vendor slides are directional.
What changed vs Eleven v3
From elevenlabs.io/v4:
- Speaker stability across regenerations — redo a line without vocal drift.
- Professional Voice Clones back with full emotional range (v3 gap).
- Multi-speaker and SFX in-script; tag following more reliable than v3.
- IPA / pronunciation dictionary for names and acronyms.
- SSML
<break>disabled — use[pause]/[long pause]tags instead.
API quick start
import { ElevenLabsClient, play } from '@elevenlabs/elevenlabs-js';
const elevenlabs = new ElevenLabsClient({
apiKey: process.env.ELEVENLABS_API_KEY,
});
const audio = await elevenlabs.textToSpeech.convert('VOICE_ID', {
text: 'The first move is what sets everything in motion.',
modelId: 'eleven_v4',
});
await play(audio);
For agents, ElevenLabs pushes Turbo with bidirectional streaming so audio begins before the LLM finishes the sentence — the same product lane as GPT-Live-1 and speech-to-speech agent guides explainx.ai already tracks.
Enterprise quotes on the v4 page
Launch social proof includes Salesforce Agentforce Voice (Ryan Peterson on deterministic control + high-EQ models) and BeyondWords (publisher engagement) — signals ElevenLabs is selling brand-safe expressive TTS into contact center and media simultaneously, not only creator tools.
What builders should test this week
- Latency — A/B v4 Turbo vs your current TTS on 200-token agent replies with your network path.
- Tags — One script with
[whispers]/[excited]/ SFX; confirm v4 follows sequence without v3-style drops. - PVC — Retrain pre-v4 clones if ElevenLabs docs require + on library voices for v4 compatibility.
- Cost — Promo $/M characters ends after two weeks — capture baseline spend before revert.
How to evaluate “#1 on Artificial Analysis”
ElevenLabs put Artificial Analysis on the launch tweet. Rankings in TTS move with voice, language, sample rate, and prompt. A #1 on a vendor-selected slice is marketing, not a bake-off you can paste into an RFP.
Run this instead:
- Same script, same voice ID, same region — v4 vs Turbo vs your current model.
- Time to first byte on a 40-token confirmation and a 200-token explanation.
- Tag fidelity —
[whispers]then[excited]in one generation; count dropped tags. - Clone drift — regenerate the same line 10 times; listen for identity hop (v4’s stated fix vs v3).
- Agent loop — stream LLM tokens into Turbo while measuring barge-in and hold-music quality, not MOS in isolation.
Salesforce Agentforce Voice’s quote on the v4 page is about deterministic agent actions plus high-EQ speech. That is the enterprise pitch: don’t improvise the tool call; do improvise the delivery. Pair with ElevenLabs Reception if you are shopping SMB phone agents, and with GPT-Live-1 if you are already on OpenAI’s realtime stack.
Privacy and clone consent
ElevenLabs states SOC 2 Type II, ISO 27001, PCI DSS Level 1, GDPR, HIPAA-eligible workflows, and verified consent for clones. Zero Retention Mode is enterprise-optional. Generated audio is meant to be detectable via their AI Speech Classifier. None of that replaces your DPA. If you clone a founder voice, keep the consent artifact.
Character cap is 10,000 per generation; long-form uses context stitching. Retrain pre-v4 IVCs/PVCs via the library + control if docs require it.
Inline tags vs SSML — copy this into your prompt library
v4 disables SSML <break>. Use natural-language tags instead:
[warm] Welcome back to the final round. [long pause]
[whispered] Beethoven. It's Beethoven, isn't it? [nervous laugh]
[delighted] Beethoven is correct! [crowd applause]
ElevenLabs’ second launch clip is the director’s-chair demo — delivery, emotion, pacing, reactions, SFX, style — without a second recording session. If tags drop, you are still on v3 habits (over-long tag stacks, SSML leftovers). Keep a pronunciation dictionary for names (Hülkenberg, Reykjavik, YAML) rather than hoping the model guesses.
v4 vs Turbo in one sentence: produced content and audiobooks stay on eleven_v4; phone agents and ElevenAgents stay on Turbo (~100 ms median inference, ~150 ms time to first speech, vendor). Promo window: $22 / $11 per 1M characters for two weeks, then standard credits (free tier 10,000 credits/month ≈ 10 minutes).
Related reading
- ElevenLabs Reception AI receptionist for SMBs
- VoiceStudio — open-source ElevenLabs-style stack
- GPT-Live-1 API for OpenAI voice agents
- Primary: elevenlabs.io/v4 · @ElevenLabs launch on X
Latency, pricing promos, and Artificial Analysis rank reflect ElevenLabs’ September 28, 2026 launch materials — verify current API pricing and model IDs in ElevenLabs docs before production.
