Skip to content
Cartesia logo

Cartesia

Low-latency voice models for real-time agents: Sonic-3.6 text to speech and Ink-2 streaming transcription, deployable in cloud, VPC or on device.

4/5 my assessment
Freemium
Artificial Intelligence

Overview

Cartesia came out of the Stanford work behind state space models, and that architecture choice is the whole product thesis. Transformers are expensive to run at low latency for streaming audio. SSMs cost less per unit of audio and degrade more gracefully on small hardware, which is how Sonic reaches roughly 90ms time to first audio under good conditions and how the same model family ends up running on a laptop. The lineup is Sonic-3.6 for text to speech, Ink-2 for streaming transcription and Line for agent orchestration on top of both.

Sonic-3.6 covers 44 languages. French and German are strong. Mandarin is competent for conversational text but lacks the tonal control you want for anything performative, so a Mandarin audiobook remains an ElevenLabs job. Pricing on the subscription tiers is credit-based rather than per character, roughly one credit per character: free at 20,000 credits a month with no commercial rights, $5 for Pro at 100,000, $49 for Startup at 1.25M and $299 for Scale at 8M. Agent minutes bill separately at $0.06, which is easy to miss when you model costs.

Anyone building a real-time voice agent should put this head to head with ElevenLabs v3 Conversational, and the deciding factor is usually deployment rather than sound quality, because both clear the bar now. If audio has to stay on a device or inside a VPC, Cartesia is one of very few options that will take the call. If you want a large stock voice library, dubbing tooling or a mature integrations ecosystem, ElevenLabs is further ahead and it is not close. Worth watching whether the on-device story extends past Apple silicon, since that is what decides if it becomes useful on embedded Linux hardware in a plant.

Key Features

  • Sonic-3.6 text to speech across 44 languages, generally available 27 August 2026, with roughly 90ms time to first audio in ideal conditions
  • Ink-2 streaming speech to text at a claimed 100ms transcript latency with native end-of-turn detection
  • The same models deployable three ways: cloud regional endpoints, VPC or on-premise, and on device
  • Edge, an Apache 2.0 open-source library with custom Metal kernels for Mamba-2, built for running state space model inference on Apple M-series silicon
  • Line, the voice agent platform, at $0.06 per minute of call plus optional telephony
  • Instant voice cloning from the $5 Pro tier, professional cloning from Startup upward

Where it holds

  • Took the #1 spot on both the Provider Voice and Controlled Voice Artificial Analysis Speech Arena leaderboards in August 2026, past Speechify's Simba 3.2 and Alibaba's Qwen-Audio-3.0-TTS-Plus
  • The state space model architecture is not a marketing line, it is the reason the same model family runs on an M-series laptop instead of requiring a datacentre GPU
  • On-device and VPC deployment in a category where nearly every competitor is cloud-only
  • At $5 a month for the Pro tier, the cost of running a proper evaluation is close to nothing

Where it breaks

  • The 90ms headline is a lab figure. Cartesia's own changelog puts real-world median time to first audio nearer 166ms to 190ms once network is in the path
  • Free tier is 20,000 credits a month and explicitly non-commercial, so any pilot that ships needs a paid plan on day one
  • Stock voice library and integration ecosystem are far thinner than ElevenLabs, which matters if you want an off-the-shelf voice rather than a cloned one
  • Ink-2 shipped in June 2026 and has no multi-vendor independent benchmark history yet, so treat the 100ms claim as something to test rather than something to plan around

My Take

Built for one job: voice agents where the gap between a human finishing a sentence and the machine starting one decides whether the call feels broken. Sonic-3.6 shipped 27 August 2026 and took the top spot on both Artificial Analysis Speech Arena boards, ahead of Speechify's Simba 3.2 and Alibaba's Qwen-Audio-3.0-TTS-Plus. The more interesting part for anyone in robotics or factory automation is the on-device path: the Edge library runs Mamba-2 kernels on M-series silicon, so a voice interface on a machine with no dependable network stops being a research project. Test the latency yourself though, because 90ms is best case and the company's own changelog puts the real-world median closer to 190ms.

Francis Okafor
Francis Okafor AI & Tech Lead · Engineer

Quick Info

Pricing:
freemium
Openness:
Proprietary
Starting at:
Free plan: 20,000 model credits/month plus $1 prepaid agent usage, non-commercial only. Pro $5/mo (100,000 credits, commercial licence, instant voice cloning). Startup $49/mo (1.25M credits, professional voice cloning, organisations). Scale $299/mo (8M credits, priority support, high concurrency). Enterprise is custom with SSO, compliance features and volume rates. Roughly 1 credit per character. Voice agents billed separately at $0.06/min of call plus $0.014/min if you use Cartesia telephony. Annual billing saves 20%. Third-party listings put Sonic-3.6 API pricing near $49 per 1M characters.
Added:
Aug 2026
Updated:
Aug 2026

Use Cases

voice agent development automation customer support

Judge it on your own work

The notes above say where Cartesia holds and where it breaks. The fastest check is your own workload.

Visit website ↗

Alternatives to Cartesia