Do Chinese AI Models Work for African Languages? The Answer Has Never Been Published
Qwen3 claims 119 languages. Open the actual list and Swahili is the only Niger-Congo entry. Yoruba appears in the technical report once, as the language they left out of the evaluation.
Do Chinese AI models work for African languages? I have spent eight years inside China's tech sector, I read the model releases in Chinese on the day they drop, and I am Nigerian. That combination ought to make this an easy question to answer. It is not, because the answer has never been published.
Not for Yoruba. Not for Igbo. Not for Nigerian Pidgin, which is the actual working language of half the engineering standups in Lagos, has an ISO code and has a FLORES-200 test set sitting there unused.
What follows is everything that is genuinely documented, which is less than you would hope and more than nothing, plus the exact procedure to produce the missing table yourself. I am going to be strict about the line between what has been measured and what has merely been asserted. That line is the whole story.
Qwen3 lists 119 languages and one of them is Niger-Congo
Qwen3 was pre-trained on roughly 36 trillion tokens spanning 119 languages and dialects, up from 29 in Qwen2.5. That headline travelled a long way. The list behind it travelled almost nowhere, because almost nobody opens it.
Open it. Alibaba groups the 119 by family on the Qwen3 release page. The Afro-Asiatic row holds eight varieties of Arabic plus Hebrew and Maltese. Hausa is an Afro-Asiatic language with one of the largest speaker populations on the continent and it is not in that row. Neither is Amharic. Neither is Somali or Oromo. The catch-all Other row is where Swahili sits, alongside Georgian, Basque, Haitian, Papiamento, Kabuverdianu and Tok Pisin. Afrikaans is filed under Indo-European, correctly.
Search that page for Hausa, Yoruba, Igbo, Amharic, Somali, Oromo, Zulu or Xhosa. None of the eight appears. Swahili is the only Niger-Congo language named among the 119. Niger-Congo is the family that holds Yoruba, Igbo, Zulu, Xhosa, Shona, Wolof and well over a hundred million Nigerian speakers on its own.
The technical report is more revealing than the blog. Qwen assessed general knowledge on MMMLU across 14 languages, "excluding the unoptimized Yoruba language" (Qwen3 Technical Report). That is the one direct quotation in this article and I am spending it here on purpose. Yoruba appears in the Qwen3 technical report exactly once, and it appears there to explain its own absence.
The Belebele result works the same way. "Belebele is a reading-comprehension benchmark in 122 language variants, a couple of dozen of them African." "Qwen reported on 80 and dropped 42 as unoptimised. Table 36 of the report lists the 80 it kept, so the 42 it dropped can be recovered exactly by subtracting that list from Belebele's 122 variants. Do the subtraction and the picture is worse than the blog suggests: the only sub-Saharan African variant Qwen kept is Swahili."

Yoruba appears in the Qwen3 technical report exactly once, and it appears there to explain its own absence.
DeepSeek, GLM and Kimi publish less than Qwen, not more
DeepSeek-V3 was trained on 14.8 trillion tokens with multilingual coverage described as expanded beyond English and Chinese, and with English and Chinese making up the majority of the corpus. There is no published language list at all. The tokenizer is a 128K byte-level BPE built for multilingual compression, which is a design intention rather than a measurement.
DeepSeek-R1's limitations section is the most useful paragraph any Chinese lab has written on this subject, precisely because it is candid. The team states that the model is optimised for Chinese and English and that queries in other languages can trigger language mixing, including reasoning and responding in English when the prompt was not English. If you have prompted R1 in Pidgin and watched it slide into English mid-answer, that behaviour is documented, expected and acknowledged by the people who shipped it.
Z.ai's GLM-4.5 sources its multilingual pre-training from its own crawl plus FineWeb-2, with a quality classifier upsampling documents judged educationally useful. FineWeb-2 does carry African languages. It also covers well over a thousand language-script pairs of which only a couple of hundred clear ten thousand documents after filtering. Upsampling a thin bucket does not thicken it. "GLM's own human evaluation uses 660 real-user prompts split into 392 English, 108 Chinese and 160 in other languages. Which other languages, the report does not say, and no African language is named anywhere in it." No African language is named anywhere in the report.
"Moonshot's Kimi K2 reports 15.5 trillion training tokens and a 163,840-token vocabulary, the largest of the Chinese flagships though still well short of Gemma 3's 262k, which is genuinely promising for morphologically dense languages.", which is genuinely promising for morphologically dense languages. Moonshot publishes no language list and no non-Chinese, non-English evaluation to sit beside it.
Not one of these labs publishes a per-language token share. Qwen 3 had the worst African-benchmark baseline of the three and it produced the largest gains after adaptation. It has to be measured.
One group measured it, and Qwen3-14B scored 40.27
The measurement exists in one serious place: the AfriqueLLM work out of McGill NLP, which continued pre-training Llama 3.1, Qwen 3 and Gemma 3 on African language data and had to establish baselines before it could claim anything. Those baselines are the numbers everyone actually wanted.
Qwen3-14B, before any adaptation, across the African benchmark suite. AfriMGSM, grade-school maths: 16.60. AfriMMLU, four-option knowledge questions: 39.66. AfriXNLI, three-way inference: 43.22. Belebele, four-option reading comprehension: 50.74. FLORES translation, English into the target language: 23.61. Injongo, intent and slot filling: 41.80. SIB-200 topic classification: 66.29. Aggregate 40.27.
Now set the chance baselines beside them, because a score without its floor is decoration. AfriMMLU has four options, so guessing scores 25. Belebele has four options, so guessing scores 25. AfriXNLI has three classes, so guessing scores 33.3. Qwen3-14B on African-language inference is about ten points above a coin toss. On African-language knowledge it is about fifteen above. On maths word problems it sits at 16.6, against an English ceiling far north of 80.
Two caveats nobody should skip. "These are means over the AfriqueLLM evaluation set, whose seven benchmarks do not all cover the same languages, and Swahili almost certainly drags the mean upward, because Swahili is the one African language on Qwen's own list.", because Swahili is the one African language on Qwen's own list. Per-language figures for Yoruba, Hausa and Igbo are not broken out in a form a builder can act on. And the suite covers Hausa, Igbo and Yoruba but not Nigerian Pidgin, Tiv, Efik, Ibibio, Kanuri or Fulfulde. Nigeria has roughly five hundred languages. The best benchmark available tests three of them.
Tokenizer fertility is the part you can check without asking permission
Fertility is tokens per word. It is the cleanest proxy for how well a model's vocabulary fits your language, it is fully determined by public tokenizer files, and it is the only number in this article a reader can regenerate before lunch.
The published figures come from a 2026 preprint measuring 13 tokenizers against FLORES-200+. English runs about 1.22 tokens per word. Yoruba comes in at 2.26 on OpenAI's o200k_base, Igbo at 1.73, Hausa at 1.64, Swahili at 1.87. Amharic, in Ethiopic script, reaches 8.97. The median across nineteen African languages is 2.29, roughly 1.88 times English.
Ranked by mean African premium, the same work places Qwen 3's 152k vocabulary fourth of eleven at about 2.63 times English, ahead of o200k_base at 2.70 but behind BLOOM's 250k vocabulary at 2.59, and places DeepSeek-V3's 129k vocabulary eighth at about 3.04. Qwen's tokenizer is comparatively kind to African orthographies. "DeepSeek's is not, and the study puts most of the blame on vocabulary size while noting that incidental multilingual coverage matters too."
This is more than a billing complaint. The Token Tax paper presented at AfricaNLP 2026 regressed AfriMMLU accuracy against fertility across ten models including DeepSeek V3 and Qwen 2.5 32B, and found consistently negative slopes in the region of eight to eighteen accuracy points lost per additional token per word, with fertility explaining a large share of the variance. Fragmentation costs money. It also predicts errors.
I want to be exact about my own position. I did not run these tokenizers. I am reporting a preprint's table, which is precisely the sort of thing that should be reproduced rather than cited, and the next section is how to reproduce it.
The evaluation that is missing costs an afternoon and a few hundred dollars
Step one, fertility. Pull the FLORES-200 dev set for eng_Latn, yor_Latn, hau_Latn, ibo_Latn and pcm_Latn. Nigerian Pidgin is in there as pcm_Latn, which means it is testable today despite appearing in no vendor's documentation. Load each tokenizer from its public repository with AutoTokenizer.from_pretrained: Qwen/Qwen3-8B, deepseek-ai/DeepSeek-V3, zai-org/GLM-4.5 and moonshotai/Kimi-K2-Instruct. For every sentence divide the token count by the word count using Unicode UAX-29 segmentation rather than a naive whitespace split, because whitespace splitting quietly flatters agglutinative languages. Report per language and divide through by your English figure. You now hold a fertility table for four models nobody has published one for.
Step two, accuracy. The IrokoBench tasks (AfriMMLU, AfriXNLI, AfriMGSM) and Belebele are merged into EleutherAI's lm-evaluation-harness, so the prompt templates and scoring are fixed and someone else's judgement calls are already settled. Point it at each model's API endpoint. Five runs, report the mean and the spread, and print the chance baseline in the same table as the score so no reader can misread a 39 as a pass.
Step three turns a benchmark into a diagnosis. Run IrokoBench's translate-test control: machine-translate the test set into English, run the model on that, then compare against native-language prompting. If translate-test wins by a wide margin the model has the reasoning and lacks the Yoruba. If both are low it lacks both. Those are different failures and they need different fixes.
Step four, publish per language and never publish the African average. The average is where Swahili hides four other languages.
"That is the entire protocol. On my own estimate it is a day or two of one engineer's time and API credit in the low thousands rather than the low millions. Which means the reason this table does not exist is unlikely to be cost."
Chinese AI models work for African languages as substrate, not as a finished product
Here is the case against my own article, and it is strong.
Published language coverage turns out to be a poor predictor of usefulness. The AfriqueLLM team compared three base families for continued pre-training on African data. Qwen 3 had the worst African-benchmark baseline of the three and it produced the largest gains after adaptation, on the order of +76.5 percent relative improvement for the 8B and +57.9 percent for the 14B. The authors call it a zero-to-hero effect and conclude that a base model's general capability is a more productive starting point than its language coverage. AfriqueQwen-14B, after 26 billion tokens across 20 African languages, moves from 40.27 to 63.58.
The field agrees by revealed preference. Sunbird AI in Kampala built Sunflower on Qwen 3 at 14B and 32B, covering 31 Ugandan languages across Bantu, Nilotic and Central Sudanic families. "Sunflower posted the best score in 24 of those 31 languages on mean bidirectional chrF, and Sunflower-32B reached a mean chrF of 0.435 translating into English, above Gemini 2.5 Pro at 0.408 and GPT-4o at 0.354.", above Gemini 2.5 Pro at 0.408 and GPT-4o at 0.354. A Ugandan team on Alibaba's open weights beat two frontier labs at Ugandan languages.
So one honest answer to the question is yes, decisively, with a condition attached. Apache-licensed weights you can continue pre-training beat a longer language list you are not permitted to touch. That advantage is real and it is why teams in Nairobi, Kampala and Lagos keep reaching for Qwen rather than waiting on a vendor roadmap.
My complaint survives in a narrower form. Adaptation is the answer for teams with GPUs and a corpus. It is not the answer for the much larger group who will simply call an API in Igbo and act on whatever comes back. For them the question is not which base to fine-tune. It is whether the raw model is right often enough to put in front of a paying customer, and that question has no published answer for any Chinese flagship in any Nigerian language.
The missing table is the opportunity
This is not a research gap. Research gaps are slow and expensive. This is a publishing gap and it is cheap.
Whoever runs the protocol above against Qwen3.5, DeepSeek V3.2, GLM and Kimi on Yoruba, Hausa, Igbo and Nigerian Pidgin, and publishes per-language numbers with chance baselines and a translate-test control, produces the reference that procurement officers, ministry advisers and startup founders will cite for the next two years. "There is no incumbent to displace, at least none I could find searching in August 2026."
One thing to carry while you build it. A model that has never been evaluated in your language is not unsupported. It is unaudited. Those are different problems and only one of them belongs to somebody else.
Sources
Qwen3 Technical Report (arXiv:2505.09388) 36T tokens, 119 languages, MMMLU excluding unoptimised Yoruba, Belebele 80 of 122 with 42 dropped: https://arxiv.org/abs/2505.09388
Qwen3 release blog. The 119-language table grouped by family: https://qwenlm.github.io/blog/qwen3/
DeepSeek-R1 paper (arXiv:2501.12948) Limitations section on optimisation for Chinese and English and language mixing: https://arxiv.org/abs/2501.12948
IrokoBench: A New Benchmark for African Languages (NAACL 2025) AfriMMLU, AfriXNLI, AfriMGSM across 17 languages: https://aclanthology.org/2025.naacl-long.139/
AfriqueLLM: How Data Mixing and Model Architecture Impact Continued Pre-training for African Languages (arXiv:2601.06395): https://arxiv.org/abs/2601.06395
AfriqueQwen-14B model card. Qwen3-14B baseline 40.27 to 63.58 after 26B tokens across 20 African languages: https://huggingface.co/McGill-NLP/AfriqueQwen-14B
The Token Tax: Systematic Bias in Multilingual Tokenization (AfricaNLP 2026) Fertility predicts AfriMMLU accuracy across 10 models: https://aclanthology.org/2026.africanlp-main.10/
The African Language Tax: Quantifying the Cost, Latency and Context Penalty of Tokenizing African Languages (arXiv:2606.24460): https://arxiv.org/abs/2606.24460
Frequently Asked Questions
Do Chinese AI models work for African languages?
There is no published per-language evaluation of Qwen, DeepSeek, GLM or Kimi on Nigerian languages, so the honest answer is that nobody has measured it properly. Qwen3-14B without adaptation scores 39.66 on AfriMMLU, where random guessing scores 25, and 43.22 on three-way inference where guessing scores 33.3. That is above chance but far below usable. After continued pre-training on a 26 billion token mixture, of which about 22.8 billion is African monolingual text and the rest is code, mathematics and synthetic translation, the same model reaches 63.58 on the aggregate suite. Chinese open-weight models work well as a base to adapt. They have not been shown to work off the shelf.
Does Qwen3 support Yoruba, Hausa or Igbo?
Not according to its own documentation. Qwen3 advertises 119 languages and dialects, but the published language table groups them by family and none of Yoruba, Hausa, Igbo, Amharic, Somali, Oromo, Zulu or Xhosa appears in it. Swahili is the only Niger-Congo language named. Afrikaans is listed under Indo-European and Kabuverdianu under the catch-all Other category. The Qwen3 technical report mentions Yoruba exactly once, to note that it was excluded from the MMMLU evaluation as an unoptimised language.
Is DeepSeek good at African languages?
DeepSeek publishes no language list, and DeepSeek-V3's technical report states that English and Chinese make up the majority of its 14.8 trillion token corpus. The DeepSeek-R1 paper is explicit in its limitations section that the model is optimised for Chinese and English and that non-English, non-Chinese queries can cause language mixing, including reasoning and answering in English regardless of the input language. On tokenizer efficiency DeepSeek-V3's 129K vocabulary fragments African languages more than Qwen3's 152K vocabulary, at roughly 3.04 times English tokens per word against Qwen's 2.63.
Which open model is the best base for building an African language model?
On the published evidence, Qwen 3. The AfriqueLLM study compared Llama 3.1, Qwen 3 and Gemma 3 as bases for continued pre-training on African languages and found Qwen 3 produced the largest gains, roughly +76.5 percent relative improvement at 8B and +57.9 percent at 14B, against +18.8 percent for Gemma 3 4B. The authors conclude that a base model's general capability matters more than its advertised language coverage. Sunbird AI's Sunflower, built on Qwen 3 at 32B, covers 31 Ugandan languages and beat Gemini 2.5 Pro and GPT-4o on mean chrF for translation into English.
What is tokenizer fertility and why does it matter for African languages?
Fertility is the average number of tokens a model's tokenizer produces per word. English typically runs about 1.22 tokens per word. Yoruba runs about 2.26 on OpenAI's o200k_base, Amharic about 8.97, and the median across nineteen African languages is 2.29. High fertility means you pay more per request, fit fewer words in the context window and wait longer for generation. It also predicts accuracy: the Token Tax paper at AfricaNLP 2026 regressed AfriMMLU scores against fertility across ten models and found roughly eight to eighteen accuracy points lost for each additional token per word.
How can I test a model on my own African language?
Two steps, both cheap. For tokenizer fertility, load the FLORES-200 dev set for your language code and English, run each public tokenizer over both with AutoTokenizer.from_pretrained, and divide tokens by words using Unicode UAX-29 segmentation. Nigerian Pidgin is available as pcm_Latn. For accuracy, the IrokoBench tasks (AfriMMLU, AfriXNLI, AfriMGSM) and Belebele are already merged into EleutherAI's lm-evaluation-harness, so you can point it at any API endpoint. Always report the chance baseline alongside the score, run IrokoBench's translate-test control to separate missing language ability from missing reasoning, and publish per language rather than as an African average.
Read next
China's University Major Cuts Are AI Policy, and Nigeria Should Read the Fine Print
The latest analysis essay.
Working on something in this space?
If this analysis is close to a problem you're thinking about, say so. I read every message personally.
Start a conversation