Skip to content
ElevenLabs logo

ElevenLabs

Text to speech, voice cloning and transcription API covering 74 languages, aimed at teams shipping narration, dubbing and voice agents.

4.5/5 my assessment
Freemium
Artificial Intelligence

Overview

The company began as a dubbing and audiobook narration tool and has grown into a full voice stack. Text to speech is still the core: Eleven v3 for expressive long-form work at $0.10 per 1,000 characters, v3 Conversational at $0.05 with roughly 280ms to first audio, and Flash or Turbo at $0.05 when you need the 75ms budget a real-time agent loop demands. Around that sit Scribe v2 for transcription at $0.22 per hour, a dubbing pipeline and an agents product billed separately at $0.08 per minute of speech.

Language coverage is the differentiator for anyone working across markets. v3 handles 74 languages, and the part that matters operationally is that a voice cloned from English source audio speaks French, German and Mandarin without a fresh clone. Quality is not uniform across all 74. European languages and Mandarin are strong and the long tail thins out fast. On the recognition side, Scribe v2 Realtime posted the lowest word error rate of any low-latency ASR model on the 30-language FLEURS benchmark, so it is competitive even in the places where it is not first.

Teams shipping narration, localised product video, accessibility audio or a customer-facing voice agent will get their money back on breadth alone. Skip it if you need on-premise or air-gapped deployment, or if the workload is high-volume streaming transcription, where Deepgram at roughly $0.29 per hour undercuts Scribe v2 Realtime at $0.39. Voice cloning also carries a consent and likeness problem that stays yours to manage, not the vendor's, and that is worth a policy decision before anyone clones a colleague.

Key Features

  • Eleven v3 text to speech across 74 languages with audio tags for emotional and pacing control, generally available since March 2026
  • v3 Conversational at roughly 280ms time to first audio, Flash and Turbo at roughly 75ms for real-time agent loops
  • Scribe v2 transcription with claimed 98%+ accuracy, plus Scribe v2 Realtime at roughly 150ms latency
  • Instant and professional voice cloning, where a voice cloned from English source audio speaks the other supported languages without re-cloning
  • Dubbing pipeline that preserves speaker identity across a language swap
  • Agents platform with keyterm prompting and entity detection as priced add-ons on the STT side

Where it holds

  • Mandarin, French and German output holds prosody under long-form narration, which is exactly where cheaper TTS starts to drift
  • The June 2026 API price cut of up to 55% changed the cost case materially, Flash and Turbo now land at $0.05 per 1,000 characters
  • One vendor covering TTS, STT, dubbing and agents cuts a lot of integration surface for a small team

Where it breaks

  • Scribe v2 loses to AssemblyAI Universal-3.5 Pro on code-switching, 8.77% WER against 7.69%, which bites if your audio mixes English and Mandarin inside a sentence
  • Diarization is a real weak spot at 35.26% on the same benchmark where AssemblyAI scores 30.17%
  • No self-hosted or air-gapped deployment, so regulated audio and plant-network workloads are ruled out
  • Credit-based subscription tiers are hard to forecast and behave differently from the flat per-character API rate, so most teams end up on two billing models at once

My Take

Two things changed in 2026 that make this worth a second look if you passed on it before: v3 went generally available in March and the June API price cut brought Flash and Turbo down to $0.05 per 1,000 characters. Multilingual is the real reason to pick it. French and German come out clean, Mandarin tone handling beats what most Western vendors manage, and a cloned voice carries across all of them without a second recording session. The soft spot is recognition, not generation: Scribe v2 sits behind AssemblyAI on code-switching and well behind on diarization, so if the recording is four people talking over each other in two languages, transcribe elsewhere and keep ElevenLabs for the voice.

Francis Okafor
Francis Okafor AI & Tech Lead · Engineer

Quick Info

Pricing:
freemium
Openness:
Proprietary
Starting at:
Free tier available. Subscriptions: Starter $5/mo, Creator $22/mo (121,000 credits, roughly two hours of speech), Pro $99/mo, Scale $330/mo. API billed in USD, not credits: Eleven v3 and v2 Multilingual TTS at $0.10 per 1,000 characters, v3 Conversational and Flash/Turbo at $0.05 per 1,000 characters. Scribe v2 speech to text $0.22/hr, Scribe v2 Realtime $0.39/hr, keyterm prompting +$0.05/hr, entity detection +$0.07/hr. Agents speech engine $0.08/min. A June 2026 update cut API rates by up to 55% and added pay-as-you-go.
Added:
Aug 2026
Updated:
Aug 2026

Use Cases

voice media creation content generation accessibility

Judge it on your own work

The notes above say where ElevenLabs holds and where it breaks. The fastest check is your own workload.

Visit website ↗

Alternatives to ElevenLabs