Skip to content

Comparison

CosyVoice vs IndexTTS vs VoxCPM: licence, cloning and cost

Three open Chinese TTS models, one decisive axis. The licence file governs whether you can ship, and for IndexTTS the Chinese text sets a revenue threshold ten times lower than the English version everyone quotes.

Last reviewed · 3 tools · 8 criteria

Verdict

Take VoxCPM2 unless you have a specific reason not to. Apache 2.0 on both weights and code, 30 languages, 48kHz output and the strongest English and Chinese seed-tts-eval figures of the three as each project reports them. Take CosyVoice instead if your work is Chinese dialect heavy or you need streaming latency, and take IndexTTS if you need explicit duration control for dubbing and your company sits under the RMB 100 million annual revenue line in the Chinese licence text. Do not take IndexTTS on the assumption that it is Apache 2.0, because it has not been for some time.

Side by side

Criterion CosyVoice IndexTTS VoxCPM
Licence on the released weights Apache 2.0. The repository LICENSE is unmodified Apache 2.0 text, and Fun-CosyVoice3-0.5B-2512 carries the apache-2.0 tag on Hugging Face.bilibili Model Use License Agreement. Not Apache 2.0 and not OSI approved. Governed by PRC law with arbitration at the Shanghai Arbitration Commission.Apache 2.0 on weights and code. The model card states it explicitly as free for commercial use.
Permission needed to ship commercially No. Apache 2.0 grants commercial use outright.Only above a size threshold. Section 2.2 requires a separate licence if you or an affiliate exceeded 100 million monthly active users last month, or crossed the annual revenue figure below.No.
The revenue threshold, English text versus Chinese text Not applicable.English LICENSE says RMB 1 billion. LICENSE_ZH.txt says 年收入超过1亿人民币, RMB 100 million. Section 9 of the English text says the Chinese version prevails, so the operative figure is the lower one.Not applicable.
Released checkpoint versus the model in the paper Paper describes a 1.5B model. Released weights are Fun-CosyVoice3-0.5B-2512 at 0.5B, in base and RL variants.IndexTTS-2.5 released with a 0.8B GPT backbone. VoxCPM's comparison table lists the earlier IndexTTS2 at 1.5B.VoxCPM2 released at 2B.
Zero-shot accuracy, as each project reports it Fun-CosyVoice3-0.5B-2512_RL card: 0.81% CER on test-zh at 77.4 similarity, 1.68% WER on test-en at 69.5, 5.44% CER on test-hard.README reports CV3-Eval, not seed-tts-eval: 6.75 WER and 73.18 speaker similarity averaged over five languages, 6.00 and 73.63 for the RL variant.README reports seed-tts-eval: 1.84% WER on test-EN at 75.3 similarity, 0.97% CER on test-ZH at 79.5, 8.13% CER on test-hard.
Languages and dialects 9 languages plus 18 or more Chinese dialects, with Guangdong, Minnan, Sichuan, Dongbei, Shanghai, Tianjin, Shandong, Ningxia and Gansu named.5: Chinese, English, Japanese, Spanish and Arabic. The last three were added in 2.5.30 languages including Arabic, Hindi, Swahili, Tagalog, Thai and Khmer, plus 9 Chinese dialects.
African-accented English Not documented. CV3-Eval's only accent subset is Chinese accent voice cloning.Not documented.Not documented. Swahili is supported as a language, which is a separate matter from African-accented English.
Speed and hardware documented Bi-streaming latency claimed as low as 150ms. Reference hardware and VRAM not documented.RTF 0.2065 at bf16 on an RTX 4090. VRAM not documented.RTF about 0.30 on an RTX 4090, about 0.13 with Nano-vLLM. Roughly 8GB VRAM, 48kHz native output, needs CUDA 12 and PyTorch 2.5 or later.

Which one to pick

  • CosyVoice

    Your work is Chinese and dialect heavy, or you are building a streaming voice agent where first-audio latency matters more than headline accuracy. Nothing else open covers 18 or more Chinese regional accents, the licence is clean Apache 2.0 and it is the smallest of the three to serve. Accept that the downloadable checkpoint is the 0.5B model, not the 1.5B one described in the paper.

  • IndexTTS

    You need exact duration control for dubbing to picture, or emotion controlled independently of timbre, and your organisation sits under 100 million monthly active users and under RMB 100 million annual revenue in the Chinese licence text. Do not pick it if you plan to distil it into another commercial model, since section 4 forbids that, or if your legal team will not accept arbitration in Shanghai under PRC law.

  • VoxCPM

    You are shipping a commercial product and want the licence question to end in one sentence, or you need language breadth beyond Chinese and English, or you want 48kHz output without upsampling. It is also the only one that designs a voice from a written description with no reference clip, which sidesteps the consent problem entirely. Budget the extra GPU: it is 2B parameters and roughly 8GB of VRAM.

I have watched this go wrong more than once in Shenzhen. A team picks a Chinese open TTS model on the strength of a demo clip, wires it into a product, then finds out at legal review that the weights were never under the licence the README appeared to promise. All three of these models are good. Only two of them can be shipped without a phone call.

CosyVoice is Alibaba's FunAudioLLM line. IndexTTS comes from bilibili. VoxCPM comes from OpenBMB, the group behind the MiniCPM models. English-language coverage of all three is thin, and where it exists it tends to copy a licence summary from someone else's blog post rather than open the file. I read the Chinese. In one case the Chinese text says something materially different from the English, and the licence itself says the Chinese wins.

The licence, and the number the English text gets wrong

Two of the three are simple. The CosyVoice repository carries the unmodified Apache License 2.0 text with no appended use restrictions, and the Fun-CosyVoice3-0.5B-2512 model card carries the apache-2.0 tag on the weights. VoxCPM is the same picture stated more bluntly: OpenBMB puts "released under the Apache-2.0 license, free for commercial use" on the model card and repeats it in the README. Both cards carry ethical language about impersonation and about demo content being illustrative. That language is not in either licence file, and the licence file is the operative grant.

IndexTTS is the one worth reading properly. In July 2025 someone opened issue #228 on the repository pointing out that the root LICENSE said Apache 2.0 while a separate INDEX_MODEL_LICENSE demanded prior written authorization for commercial use. No maintainer answered. That ambiguity is what most English write-ups froze in amber, and it is why you still see "IndexTTS requires written bilibili authorization for any commercial use" repeated as fact.

As of 1 September 2026 the repository resolves it differently. There is one LICENSE, the bilibili Model Use License Agreement, alongside LICENSE_ZH.txt and a DISCLAIMER. No Apache text remains anywhere in the tree. Section 2.1 grants a worldwide, non-exclusive, non-transferable, royalty-free limited licence to use the model. There is no blanket commercial bar. Section 2.2 instead sets a scale threshold in the Meta style: you must request a separate licence only if you or an affiliate had more than 100 million monthly active users in the preceding calendar month, or if annual revenue in the preceding year crossed a stated figure.

That stated figure is where the English and Chinese diverge. The English LICENSE says RMB 1 billion. LICENSE_ZH.txt says 年收入超过1亿人民币, which is RMB 100 million, one tenth as much. Section 9 of the English text settles which one applies: where the two versions conflict, the Chinese version prevails and governs the rights and obligations of the parties. So the real trigger is RMB 100 million, and a mid-sized Chinese media company that reads only the English text will conclude it is clear when it is not. Section 6 sends any dispute to the Shanghai Arbitration Commission under PRC law.

Two more clauses matter in practice. Section 4 bars using the model or a derivative work to improve any other AI model, with a carve-out for IndexTTS itself, its derivatives and non-commercial models, so distilling it into your own commercial voice model is out. And the DISCLAIMER's line about 未经授权将合成声音用于商业目的 is about cloning a real person's voice without that person's authorization, not about commercial use of the model. English summaries routinely conflate the two. One loose end: the licence defines the model by name as "bilibili indextts2", while the current release is IndexTTS-2.5. The 2.5 model card points at the same agreement, so the intent is clear enough, but the definition has not been updated to match.

What the papers actually report

The CosyVoice 3 paper describes a 1.5B model and reports 0.71% CER with 0.836 speaker similarity on SEED test-zh, and 1.45% WER with 0.784 on test-en, after reinforcement learning. Those weights were not released. What you can download is Fun-CosyVoice3-0.5B-2512, at 0.5B, and its card publishes its own numbers: 1.21% CER and 78.0 similarity on test-zh for the base checkpoint, 2.24% WER and 71.8 on test-en, with the RL variant at 0.81% and 1.68% respectively. Quoting the paper's figures for the checkpoint you actually ran is a common and consequential error.

VoxCPM2's README reports 1.84% WER at 75.3 similarity on seed-tts-eval test-EN, and 0.97% CER at 79.5 on test-ZH. Its comparison table places IndexTTS2 at 2.23% and 1.03%. Those IndexTTS numbers were produced by the VoxCPM team, not published by bilibili, and a vendor's table of its competitors deserves the scepticism you would apply to any vendor's table of its competitors.

IndexTTS's own reporting is where comparisons break down entirely. The IndexTTS2 paper claims to beat state of the art on word error rate, speaker similarity and emotional fidelity, without numbers in the abstract. The IndexTTS-2.5 README does not report seed-tts-eval at all. It reports CV3-Eval, Alibaba's benchmark, at 6.75 WER and 73.18 speaker similarity averaged across Chinese, English, Japanese, Spanish and Arabic, improving to 6.00 and 73.63 for the RL variant. A five-language average on a harder benchmark is not the same measurement as 1.84 on English seed-tts-eval. Anyone who puts those two figures in one table is misleading you, whether or not they mean to.

The accuracy story is also not one-sided. On seed-tts-eval's hard set, CosyVoice's RL checkpoint reports 5.44% CER against VoxCPM2's 8.13%, and VoxCPM's own table gives IndexTTS2 7.12%. And the CosyVoice card publishes a human baseline row of 1.26% CER on test-zh, which the RL checkpoint beats. That tells you the metric has saturated, not that the synthesis is better than a person. Where IndexTTS genuinely leads is control: an eight-float emotion vector disentangled from timbre, and a mode that fixes the generated token count so you can hit a target duration exactly. For dubbing to picture, that is the feature that matters, and neither of the other two offers its equivalent. The card is honest that enabling random emotion sampling degrades cloning fidelity.

Languages, dialects and the accent nobody trained on

CosyVoice covers nine languages: Chinese, English, Japanese, Korean, German, Spanish, French, Italian and Russian, plus what the card calls 18 or more Chinese dialects, with Guangdong, Minnan, Sichuan, Dongbei, Shanghai, Tianjin, Shandong, Ningxia and Gansu named explicitly. It was trained on a million hours. Nothing else in the open field comes close on Chinese regional speech, and CV3-Eval's only accent subset is Chinese accent voice cloning.

IndexTTS-2.5 covers five: Chinese, English, Japanese, Spanish and Arabic, with the last three added in the 2.5 release. VoxCPM2 covers thirty, including Arabic, Hindi, Swahili, Tagalog, Vietnamese, Thai and Khmer, plus nine Chinese dialects. If your requirement is breadth, that is not a close contest.

On African-accented English, the honest answer is that none of the three documents it. CosyVoice 2's data description listed Chinese, Indian and Russian English accents as speaking styles. The CosyVoice 3 paper does not mention African accents anywhere. VoxCPM2 supports Swahili as a language, which is a different thing from Nigerian or Ghanaian English. IndexTTS says nothing on the subject.

That leaves zero-shot cloning from your own reference clip as the only route, and here is the part that should worry you: none of these projects publishes an accent-preservation metric. Speaker similarity on seed-tts-eval measures timbre against a mostly Chinese and American reference distribution. It will not tell you whether a synthesised Lagos or Nairobi voice keeps its vowel system and its even-paced rhythm, or drifts toward General American on longer passages. I have not run a controlled study, so I will not pretend to a number. What I will say is that this is the one axis you must evaluate yourself, with your own speakers and your own listeners, before you commit. Treat every published similarity figure as silent on it.

Hardware and what a second of audio costs you

All three are self-hosted, so the cost line is GPU-hours rather than per-character API billing. That changes the calculation entirely if your volume is high and your latency budget is loose.

VoxCPM2 is the heaviest at 2B parameters, roughly 8GB of VRAM and a real-time factor of about 0.30 on an RTX 4090, dropping to about 0.13 with Nano-vLLM acceleration. It needs Python 3.10 or later below 3.13, PyTorch 2.5 or later and CUDA 12. In exchange you get native 48kHz output that you do not have to upsample, plus voice design from a written description with no reference clip at all.

IndexTTS-2.5 is 0.8B and reports 0.2065 RTF at bf16 on the same RTX 4090, so roughly five times faster than real time. VRAM is not documented. CosyVoice's released checkpoint is the smallest at 0.5B and the card claims bi-streaming latency as low as 150ms, but names no reference hardware for that figure and states no VRAM requirement, so treat it as a ceiling rather than a spec. For a streaming voice agent, the streaming architecture matters more than the raw RTF, and CosyVoice is the only one of the three built around it.

When none of the three is the answer

If you are building a conversational agent that has to answer inside 200 milliseconds end to end, self-hosting any of these and then attaching ASR, an LLM and turn detection is a larger engineering programme than buying a hosted low-latency voice API. The model is the easy part of that stack.

If you need indemnification against a voice-likeness claim, none of these three gives it to you. Apache 2.0 disclaims warranties, and the bilibili agreement puts the burden of defending third-party infringement claims squarely on you. Any commercial deployment that clones identifiable people needs signed releases regardless of which model produced the audio.

And if your actual target is a low-resource African language rather than accented English, none of the three trained on it in any meaningful volume. The real project there is a fine-tune on your own corpus, which makes the licence question decisive again for a different reason: you can fine-tune and ship either Apache 2.0 model freely, while section 4 of the bilibili agreement restricts using IndexTTS or its derivatives to improve other models. That clause alone rules IndexTTS out of most serious localisation work.

Questions

Is IndexTTS Apache 2.0 or not?

Not any more, and it is worth being precise about this. Through mid-2025 the repository carried both an Apache 2.0 LICENSE and a separate INDEX_MODEL_LICENSE that demanded prior written authorization for commercial use, a contradiction raised in issue #228 in July 2025 and never answered by a maintainer. As of 1 September 2026 the repository carries a single LICENSE, the bilibili Model Use License Agreement, plus a Chinese version and a DISCLAIMER. No Apache text remains. Anything you read that calls IndexTTS Apache 2.0 is describing a state of the repository that no longer exists.

Can I use IndexTTS commercially without contacting bilibili?

Section 2.1 grants a royalty-free licence to use the model with no blanket commercial bar. Section 2.2 requires a separate licence only if you or an affiliate had more than 100 million monthly active users in the preceding month, or if annual revenue crossed the stated threshold. The English text puts that threshold at RMB 1 billion and the Chinese text at RMB 100 million, and section 9 says the Chinese version prevails, so plan against the lower figure. This is a reading of the licence text, not legal advice, and any real deployment should be reviewed by counsel who reads the Chinese.

Which one clones a Nigerian or other African-accented English voice best?

None of the three documents African-accented English, none publishes an accent-preservation metric, and I have not run a controlled listening test, so I will not give you a ranking dressed up as a finding. What I can tell you is that seed-tts-eval speaker similarity will not answer the question for you, because it measures timbre against a reference distribution that is largely Chinese and American. Zero-shot cloning from your own reference clip is the only route on all three, and the failure mode to listen for is accent drift toward General American over longer passages. Test with your own speakers before you commit.

Why do the accuracy numbers disagree so much between reviews?

Because they come from different benchmarks. CosyVoice and VoxCPM both report seed-tts-eval, so their CER and WER figures sit on the same scale. IndexTTS-2.5's README reports CV3-Eval instead, averaged across five languages, which produces a much higher error rate on a harder test. Reviews that put 1.84 and 6.00 in the same column are comparing two different measurements. On top of that, VoxCPM's table of IndexTTS results was produced by the VoxCPM team, not published by bilibili.

Does any of them come with commercial indemnification?

No. Apache 2.0 disclaims warranties, and the bilibili agreement explicitly puts the burden of defending third-party infringement claims on the user. If your risk model requires indemnification against a voice-likeness or training-data claim, a hosted commercial vendor that offers it is a different product category from these three, and the licence comparison here does not substitute for that.

Which is cheapest to run at volume?

CosyVoice's 0.5B checkpoint and IndexTTS-2.5's 0.8B model are both small enough to serve comfortably on a single consumer GPU, and IndexTTS reports 0.2065 RTF at bf16 on an RTX 4090. VoxCPM2 at 2B is the heaviest, around 8GB of VRAM and roughly 0.30 RTF on the same card, though Nano-vLLM brings that to about 0.13. Since all three are self-hosted, the meaningful cost comparison is GPU-hours against your own throughput, and none of these projects publishes hosted pricing to compare against.

Sources

  1. CosyVoice repository LICENSE (Apache 2.0) github.com
  2. Fun-CosyVoice3-0.5B-2512 model card, evaluation table and dialect list huggingface.co
  3. CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training (arXiv 2505.17589) arxiv.org
  4. bilibili Model Use License Agreement, English text github.com
  5. bilibili Model Use License Agreement, Chinese text (LICENSE_ZH.txt) github.com
  6. IndexTTS repository README, benchmarks and RTF github.com
  7. IndexTTS licence ambiguity, issue #228 github.com
  8. IndexTTS-2.5 model card huggingface.co
  9. IndexTTS2 paper (arXiv 2506.21619) arxiv.org
  10. VoxCPM repository README, benchmark table and language list github.com
  11. VoxCPM2 model card huggingface.co
  12. VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation (arXiv 2509.24650) arxiv.org

Individual reviews: CosyVoice, IndexTTS, VoxCPM. All comparisons, or the full tool catalogue.