Skip to content
💡 Innovation Ecosystem

AI in the Global South Is Priced in Dollars and Written in English

Francis Okafor Francis Okafor
10 min read
AI and society Global South Africa AI policy open weights AI labour multilingual AI
AI in the Global South Is Priced in Dollars and Written in English
On this page
  1. Ninety-eight percent of African languages are missing
  2. Twenty dollars is not twenty dollars
  3. The annotation economy pays in shillings and bills in dollars
  4. What a credential signals when everyone has the same tool
  5. The governance gap is a staffing gap
  6. Where it has genuinely levelled things
  7. The strongest case against everything above
  8. The part that does not resolve
  9. Tools referenced
  10. Sources

Walk the third floor of the SEG electronics market in Huaqiangbei and you can buy a Rockchip RK3588 module with an on-board NPU for a few hundred yuan. Cash, no invoice, no export paperwork. It will run a quantised detection model at video rate on single-digit watts. Nobody in that building has an opinion about superintelligence. They have opinions about yield, unit cost and whether the buyer in Onitsha will wire the deposit. That is roughly the distance between how AI in the Global South gets discussed and how it actually gets transacted.

The dominant conversation about AI and society is written in San Francisco, London and Brussels. It is a serious conversation and I read it closely. But it is calibrated to risks that bite hardest in high-income countries: displacement of knowledge work, synthetic media in mature democracies, long-horizon catastrophic risk. Stand in Lagos or Shenzhen and a different set of questions moves to the front. Who is in the training data. Who can pay for inference. Whose language gets served, at what price per sentence. And whether a technology built somewhere else arrives as a tool you pick up or a subscription you rent.

Ninety-eight percent of African languages are missing

A survey published in June 2025 by Kedir Yassin Hussen and five co-authors counted the African languages actually supported across the large and small language models then available. The number was 42, backed by 23 public datasets. Four appear consistently: Amharic, Swahili, Afrikaans and Malagasy. Africa has roughly 2,000 languages. More than 98% of them have no usable model support at all.

That is a coverage problem. It is also, less obviously, a pricing problem. Aleksandar Petrov, Emanuele La Malfa, Philip Torr and Adel Bibi showed at NeurIPS 2023 that the same sentence translated across languages produces tokenizer outputs differing in length by up to fifteen times. Commercial APIs bill per token. A tokenizer trained mostly on English chops Yoruba or Amharic into more and smaller pieces, so the identical question costs more to ask, fills the context window faster and comes back worse. The bias is not only in the answer. It is in the meter.

I have watched this with Igbo. Ask a frontier model something in clean Igbo and it will often reply in a register nobody in Enugu uses, having quietly translated to English internally, reasoned there and translated back. The output is grammatical. It is also foreign.

There are responses. In September 2025 Nigeria's Federal Ministry of Communications, Innovation and Digital Economy, working with NCAIR and NITDA, released N-ATLaS through the startup Awarri: four speech recognition models and an 8-billion-parameter text model on a Llama-3 base, covering Yoruba, Igbo, Hausa and Nigerian-accented English, with Pidgin added afterwards. Lelapa AI's InkubaLM is a 400-million-parameter model for five African languages, shrunk by 75% during 2025 through a competition run with Zindi. Small, unglamorous, locally owned.

How a choice made during pretraining turns into a price in Lagos. A tokenizer fitted mostly to English splits Yoruba, Igbo or Amharic into more and smaller pieces, so the same question consumes more billable tokens, exhausts the context window faster and returns a weaker answer. The penalty compounds because the bill is denominated in dollars while the income is not.
How a choice made during pretraining turns into a price in Lagos. A tokenizer fitted mostly to English splits Yoruba, Igbo or Amharic into more and smaller pieces, so the same question consumes more billable tokens, exhausts the context window faster and returns a weaker answer. The penalty compounds because the bill is denominated in dollars while the income is not.
High trust in a system you did not build, cannot audit and cannot price is a description of exposure, not of power.

Twenty dollars is not twenty dollars

Nigeria's national minimum wage is 70,000 naira a month, set by the 2024 amendment act. At the Central Bank rate in late August 2026, around 1,340 naira to the dollar, that is about $52.

OpenAI began collecting Nigeria's 7.5% VAT on 1 November 2025, which puts a ChatGPT Plus seat at $21.50 a month. Against the statutory minimum wage that is roughly 41% of gross monthly income, for one product. A cheaper Nigerian tier followed at 7,000 naira, about $5. That is a real concession and also an admission that flat dollar pricing does not survive contact with a naira income.

The hardware underneath has the same shape. GSMA's November 2025 report on smartphone adoption in Africa puts the median cost of an entry-level smartphone in Sub-Saharan Africa at about $39 in 2024, up from $38 the year before, equal to 26% of monthly GDP per capita against a 16% average across low- and middle-income countries. Africa holds 33% of the world's unconnected population, and most of those people already live inside broadband coverage. The network is there. The device is not.

None of this is malice. It is what happens when a product priced in one currency zone is offered at a single global number and called accessible.

The annotation economy pays in shillings and bills in dollars

In January 2023 TIME reported that OpenAI had used Sama, a San Francisco outsourcing firm, to label toxic content in Kenya. Workers took home between $1.32 and $2 an hour. OpenAI paid Sama $12.50 an hour for the same labour. The contracts ran to roughly $200,000, later described by OpenAI as about $150,000 after Sama ended the work eight months early.

That story is three and a half years old. The structure has not improved. It has gone darker. Rest of World reported on 4 December 2025 that Chinese AI firms are recruiting Kenyan students and recent graduates through WhatsApp groups and Google Forms, paying by M-Pesa. Workers described shifts of up to twelve hours for as little as 700 Kenyan shillings, about $5.42. Daily quotas ran into tens of thousands of video clips. Accuracy below 85% meant no payment. Most could not name the company whose model they were training.

Kenya has noticed. The draft Kenya Artificial Intelligence and Other Emerging Technologies Policy, published in July 2026 with the consultation closing on 4 August, proposes something no other jurisdiction has seriously attempted: a fair-pay reference framework requiring firms to benchmark annotator pay against international rates for equivalent work, then publish the comparison. It adds duty-of-care obligations for exposure to disturbing material, mandatory mental health provision, written contracts and grievance channels.

Whether it becomes law is a separate matter. The drafting is the interesting part. It is the first time the input side of this industry has been priced as labour rather than as compute.

What a credential signals when everyone has the same tool

Nigeria's 3 Million Technical Talent programme was built to train developers. For N-ATLaS it doubled as a mobilisation channel, recruiting ordinary Nigerians to record and transcribe speech in their own languages. A state skills pipeline operating as a data-collection workforce. It was a fair trade in that instance, and it is also a preview of what education systems here are being asked to become.

The credential problem bites harder outside rich countries because the credential carries more weight. Where hiring is remote, references are thin and the employer is three time zones away, a certificate and a timed coding test are most of the signal. Both are now trivially defeated by tools costing five dollars a month. Institutions in Lagos, Nairobi and Accra are being asked to defend assessment integrity on budgets that were already short for lab equipment. Universities that ban the tools produce graduates who cannot use them. Universities that ignore the problem produce transcripts nobody trusts. There is not yet a third option and the deadline is now.

The governance gap is a staffing gap

The IMF's AI Preparedness Index scores 174 economies on digital infrastructure, human capital, innovation and regulation. Advanced economies average 0.68. Emerging markets average 0.46. Low-income countries average 0.32.

Policy text is not the shortage. The African Union adopted its Continental Artificial Intelligence Strategy in July 2024 with a 2025 to 2030 implementation window. Nigeria has a national AI strategy. Kenya has a draft policy out for comment. Documents exist.

What is scarce is people who can read a model card and a procurement contract in the same afternoon. A ministry with four technically literate staff cannot audit a vendor's claims about a face recognition system, cannot check whether a fine-tune actually removed the behaviour the vendor says it removed and cannot separate a genuine safety case from a slide deck. So it signs. Regulatory capacity is a hiring line, not a legal question, and hiring lines are what finance ministries cut first.

Where it has genuinely levelled things

The open-weights turn is the largest single equaliser of the last three years, and it did not come from a development agency.

Alibaba's Qwen3 expanded multilingual coverage from 29 languages to 119 languages and dialects, released under Apache 2.0 in sizes from 0.6 billion to 235 billion parameters. Google's Gemma line and Mistral's releases sit in the same category. Practically, that means a team in Kano can pull the weights, fine-tune on their own corpus and change the tokenizer behaviour that was costing them money. The per-token bill becomes an electricity bill. llama.cpp, Ollama, vLLM, ncnn and MNN exist to make that run on hardware people already own.

The price floor collapsed too. DeepSeek's published rates in August 2026 put its flash model at $0.22 per million input tokens off-peak and $0.66 per million output, with cached input at fractions of a cent. Competition from Chinese labs has done more for affordability in African deployments than any access programme I am aware of.

Compute is arriving, slowly. Cassava Technologies has deployed NVIDIA-powered capacity in South Africa and sells GPU access as a service, with announced expansion into Nigeria, Kenya, Egypt and Morocco.

And the attitude data does not read like grievance. The University of Melbourne and KPMG surveyed more than 48,000 people across 47 countries between November 2024 and January 2025. Roughly three in five respondents in emerging economies said they were willing to trust AI systems, against two in five in advanced economies. Self-reported AI literacy ran 64% against 46%. Training, 50% against 32%. Optimism was the dominant emotion in emerging economies. Worry dominated in many rich ones.

The strongest case against everything above

Here is the argument that gives me the most trouble, and it is not a weak one.

The dependency framing is self-serving and largely wrong. Nobody is compelled to rent inference. Near-frontier weights from Alibaba, DeepSeek, Google and Mistral are free to download today under permissive licences. The tokenizer penalty vanishes the moment you train your own. The binding constraints on AI in the Global South are electricity, bandwidth, capital and state capacity, and not one of those is exportable from California. The IMF's own analysis makes the point bluntly: sub-Saharan Africa could gain around 4% of GDP over a decade from AI given better power, connectivity and skills, and close to nothing without them. Blaming model providers is a way of not building substations.

The survey data cuts the same direction. The people supposedly being colonised are, measurably, among the most enthusiastic adopters on earth.

I concede most of this. The substation point especially is correct and under-argued across the continent, where it is more comfortable to write about extraction than about transmission losses and diesel.

Two things survive it.

Defaults beat options. That you can train your own tokenizer is true and irrelevant to a fifteen-year-old in Aba typing Igbo into a free chat window. She gets the default. The default charges more tokens for the same sentence and answers her in a register that is not hers. Optionality requiring a GPU, a corpus and six months is not optionality for the people actually affected.

And enthusiasm is not leverage. High trust in a system you did not build, cannot audit and cannot price is a description of exposure. The students labelling video clips through WhatsApp middlemen for 700 shillings a shift are enthusiastic adopters too. Enthusiasm was never the variable in question.

The part that does not resolve

Two images I hold at once.

The module in Huaqiangbei, sold to anyone with cash. A workshop in Aba can put a working vision system on a line for the price of a mid-range phone, inside a week, without a licence, a partner agreement or a single meeting with anyone in California. Real levelling, and it happened because Chinese hardware margins collapsed, not because anyone intended it.

And the WhatsApp group in Nairobi paying $5.42 for twelve hours, feeding the same industry from the other end. Both are AI arriving in the Global South. Both are the same supply chain.

What I cannot resolve is that the open weights making the first image possible are published by companies whose reasons have nothing to do with Nigeria or Kenya. Free is a competitive strategy. Strategies end. Everyone I know building on open models understands this and builds anyway, because the alternative is not building, and eight years here has taught me that not building is how you end up renting forever.

Nobody has told us when the strategy expires. We are all shipping against a clock we cannot see.

Tools referenced

Ollama, reviewed here: Ollama review.

llama.cpp, reviewed here: llama.cpp review.

vLLM, reviewed here: vLLM review.

ONNX Runtime, reviewed here: ONNX Runtime review.

DeepSeek, reviewed here: DeepSeek review.

Google Gemma 4, reviewed here: Google Gemma 4 review.

Sources

Hussen et al., The State of Large Language Models for African Languages (arXiv, June 2025): https://arxiv.org/abs/2506.02280

Petrov et al., Language Model Tokenizers Introduce Unfairness Between Languages (NeurIPS 2023): https://arxiv.org/abs/2305.15425

TIME: OpenAI used Kenyan workers on less than $2 per hour (January 2023): https://time.com/6247678/openai-chatgpt-kenya-workers/

Rest of World: Chinese tech companies hire Kenyan workers for AI training (December 2025): https://restofworld.org/2025/kenya-china-ai-workers/

Draft Kenya Artificial Intelligence and Other Emerging Technologies Policy, July 2026: https://ict.go.ke/sites/default/files/AI%20Policy%20Doc/draft-kenya-ai-and-emerging-technologies-policy-2026.pdf

KPMG and University of Melbourne, Trust, Attitudes and Use of AI: A Global Study 2025: https://kpmg.com/xx/en/our-insights/ai-and-technology/trust-attitudes-and-use-of-ai.html

African Union, Continental Artificial Intelligence Strategy (July 2024): https://au.int/sites/default/files/documents/44004-doc-EN-_Continental_AI_Strategy_July_2024.pdf

N-ATLaS model card, NCAIR / Awarri (Hugging Face, September 2025): https://huggingface.co/NCAIR1/N-ATLaS

Frequently Asked Questions

How many African languages do AI language models actually support?

A survey published in June 2025 by Kedir Yassin Hussen and colleagues counted 42 African languages supported across the large and small language models then available, backed by 23 public datasets. Only four appear consistently: Amharic, Swahili, Afrikaans and Malagasy. Africa has roughly 2,000 languages, so more than 98% have no usable model support. Coverage is widening through open-weights releases such as Alibaba's Qwen3, which spans 119 languages and dialects, and through local projects including Nigeria's N-ATLaS (Yoruba, Igbo, Hausa, Nigerian-accented English and Pidgin) and Lelapa AI's 400-million-parameter InkubaLM.

How much are AI data annotation workers paid?

Low, and it varies by country and task. TIME reported in January 2023 that Kenyan workers labelling toxic content for OpenAI through the outsourcing firm Sama took home $1.32 to $2 an hour, while OpenAI paid Sama $12.50 an hour for the same work. Rest of World reported in December 2025 that Kenyan students labelling video for Chinese AI firms earned as little as 700 Kenyan shillings, about $5.42, for shifts of up to twelve hours, recruited through WhatsApp groups and paid via M-Pesa. Kenya's draft July 2026 AI policy proposes forcing firms to benchmark that pay against international rates for equivalent work and publish the comparison.

Do people in developing countries trust AI more than people in rich countries?

Yes, consistently. The University of Melbourne and KPMG surveyed more than 48,000 people across 47 countries between November 2024 and January 2025. Roughly three in five respondents in emerging economies said they were willing to trust AI systems, against two in five in advanced economies. Emerging-economy respondents also reported higher AI literacy (64% against 46%) and more training (50% against 32%), and optimism was the dominant emotion there, while worry dominated in many high-income countries.