Skip to content
Deepgram logo

Deepgram

Speech recognition API for production voice pipelines: Nova-3 batch and streaming, Flux for turn-aware agents and a supported self-hosted deployment path.

3.5/5 my assessment
Paid
Artificial Intelligence

Overview

Deepgram sells speech recognition as infrastructure rather than as a product with a UI. Nova-3 is the general model, priced per minute with splits between monolingual and multilingual and between batch and streaming: $0.0043 and $0.0048 per minute respectively on pay-as-you-go for English. Flux is the newer and more interesting piece, a streaming model with turn detection built into the transcription rather than bolted on top, so a voice agent stops interrupting people mid-sentence. Aura-2 covers text to speech and a Voice Agent API at $0.056 per minute wraps the whole loop for teams that do not want to assemble it.

For anyone operating across languages the 2026 Asia-Pacific expansion is the change that matters. Nova-3 added Mandarin Simplified, Mandarin Traditional and Cantonese Traditional, with a claimed 65.21% relative word error reduction on Mandarin Simplified over Nova-2 and 24.82% on Cantonese. French and German have been solid for considerably longer. Real-time multilingual with code-switching does exist, but the independent numbers put it last among the three big vendors at mixing languages inside a single sentence, so a French engineer dropping into English technical terms mid-thought will generate errors somebody has to fix.

The right buyer has volume and a deployment constraint. Contact-centre analytics, meeting capture at scale, voice agents where per-minute cost compounds fast or any pipeline where the audio has to stay inside your own network. The wrong buyer wants one vendor for everything, because Aura-2 will not satisfy anyone who has heard Sonic-3.6 or Eleven v3, and low-volume high-accuracy work is better served by paying AssemblyAI or ElevenLabs and not thinking about it again.

Key Features

  • Nova-3 speech to text with separate batch and streaming endpoints, priced per minute from $0.0043
  • Real-time multilingual transcription with code-switching, covering Mandarin Simplified, Mandarin Traditional and Cantonese Traditional after the 2026 Asia-Pacific expansion
  • Flux, a streaming conversational STT model with end-of-turn detection built in, so an agent stops guessing when the human finished speaking
  • Keyterm prompting for product names, part numbers and jargon without training a custom model
  • Aura-2 text to speech and a Voice Agent API at $0.056 per minute that wraps the full STT to LLM to TTS loop
  • Self-hosted and VPC deployment, with documented patterns alongside AWS, NVIDIA and Fortanix

Where it holds

  • Cheapest credible streaming ASR at roughly $0.29 per hour against AssemblyAI's $0.45, and that gap compounds hard at contact-centre volume
  • Self-hosted deployment is an actual supported product rather than a sales conversation, which is the deciding factor for factory-floor and regulated audio
  • Nova-3 cut Mandarin Simplified word error rate by a claimed 65.21% against Nova-2 and Cantonese Traditional by 24.82%, so Chinese-language audio finally reaches usable

Where it breaks

  • The accuracy lead is gone: 12.22% WER on the 2026 code-switching benchmark against AssemblyAI Universal-3.5 Pro at 7.69%
  • Diarization is the weakest of the three majors at 37.92%, so multi-speaker meeting transcripts still need a cleanup pass
  • Aura-2 TTS is audibly behind Sonic-3.6 and Eleven v3 on naturalness, which pushes most teams into paying two vendors anyway
  • Growth is a committed spend near $4,000 a year, an awkward jump from pay-as-you-go for a mid-size team

My Take

The accuracy lead is gone. On the 2026 code-switching benchmark Nova-3 Multilingual posts 12.22% WER against AssemblyAI Universal-3.5 Pro at 7.69%, and diarization is worse still at 37.92%. What Deepgram kept is the pair of things that usually decide a production deployment: unit price and where the model is allowed to run. Streaming at roughly $0.29 an hour undercuts AssemblyAI by about a third, and self-hosted is a documented deployment path rather than an enterprise promise, which is the entire argument for anyone transcribing audio that cannot leave a plant network or a hospital VPC.

Francis Okafor
Francis Okafor AI & Tech Lead · Engineer

Quick Info

Pricing:
paid
Openness:
Proprietary
Starting at:
$200 free credit on signup, roughly 45,000 minutes. Pay As You Go: Nova-3 pre-recorded $0.0043/min monolingual and $0.0052/min multilingual, streaming $0.0048/min monolingual and $0.0058/min multilingual. Flux streaming $0.0065/min English, $0.0078/min multilingual. Aura-2 TTS $0.030 per 1,000 characters, Aura-1 $0.0150. Voice Agent API $0.056/min. Growth tier is a committed-spend plan around $4,000/year that discounts every rate by roughly 12 to 15%. Self-hosted and enterprise pricing is custom.
Added:
Aug 2026
Updated:
Aug 2026

Use Cases

voice agent development enterprise ai customer support

Judge it on your own work

The notes above say where Deepgram holds and where it breaks. The fastest check is your own workload.

Visit website ↗

Alternatives to Deepgram