Skip to content
💡 Innovation Ecosystem

LLM API Pricing in Dollars, Revenue in Naira

Francis Okafor Francis Okafor
8 min read
AI economics Tokenization African tech LLM APIs Unit economics Multilingual NLP
LLM API Pricing in Dollars, Revenue in Naira
On this page
  1. The dollar is the denominator
  2. What LLM API pricing actually charges you for
  3. The tokenizer charges Yoruba more than English for the same sentence
  4. Two taxes, one invoice
  5. The strongest objection: open weights make all of this temporary
  6. What actually moves the number
  7. The decision nobody thinks of as a pricing decision
  8. Tools referenced
  9. Sources

A friend in Lagos priced his product at ₦2,000 a month. Reasonable number. Roughly a mid-tier data bundle. Then he opened the LLM API pricing page for the model his product runs on and found every figure denominated in dollars per million tokens. On 28 August 2026 the CBN official rate sat near ₦1,338 to the dollar, so his ₦2,000 was $1.49. Everything after that was arithmetic he had not planned to do.

Metered intelligence turns a product decision into a foreign exchange position. You are not buying a model. You are buying a stream of tokens at a dollar price, selling something at a naira price and standing in the gap. Most writing about model costs quietly assumes those two prices live in the same currency. For a large share of the people building right now, they do not.

Two separate taxes operate here and they compound. One is the currency. The other is the tokenizer, which decides how many billable units your sentence becomes before anyone has run a single forward pass.

The dollar is the denominator

Nigeria floated the naira in June 2023. On 14 June it moved from roughly ₦470 to about ₦750 against the dollar inside a single session, a fall of around 36 percent. By early 2024 it had reached ₦1,600. It has since settled into a band: the CBN official close was ₦1,346.90 on 25 August 2026, with the parallel market quoting nearer ₦1,400.

Sit with what that did to a subscription. A product priced at ₦2,000 in May 2023 was worth about $4.30. The same ₦2,000 today is $1.49. The founder did not cut prices. The founder did not lose a single customer. The dollar cost of every API call they make simply tripled relative to what they collect.

Nigeria's minimum wage is ₦70,000 a month. When it was agreed in 2024 that was around $50. Today it is closer to $42. That is the ceiling consumer pricing has to fit under, and it is denominated in the currency that keeps moving the wrong way.

Tokenizer fertility on o200k_base measured over the FLORES-200+ parallel corpus, where every language is expressing the identical sentence. A 20-word reply is about 24 tokens in English and about 217 in N'Ko. Because the price per token is the same for every language, the same conversation bills nine times over.
Tokenizer fertility on o200k_base measured over the FLORES-200+ parallel corpus, where every language is expressing the identical sentence. A 20-word reply is about 24 tokens in English and about 217 in N'Ko. Because the price per token is the same for every language, the same conversation bills nine times over.
There is no naira sheet. There is no cedi sheet. Every African builder does that conversion in their head, every time.

What LLM API pricing actually charges you for

Here is the sheet on 29 August 2026, taken from the vendors' own documentation rather than the comparison blogs, which are wrong more often than they are right.

Anthropic lists Claude Opus 5 at $5 per million input tokens and $25 output; Claude Sonnet 5 at $2 and $10; Claude Haiku 4.5 at $1 and $5; Claude Fable 5 at $10 and $50. OpenAI lists gpt-5.6-sol at $4 and $20; gpt-5.6-terra at $2 and $12; gpt-5.6-luna at $0.20 and $1.20; gpt-5-nano at $0.05 and $0.40. Google lists Gemini 3.7 Flash at $0.75 and $3.75 through 31 December 2026, after which it doubles to $1.50 and $7.50. DeepSeek lists v4-flash at $0.22 input and $0.66 output off-peak, doubling to $0.44 and $1.32 at peak.

The cheapest verified output token there is gpt-5-nano at $0.40 per million. The most expensive is Fable 5 at $50. A spread of 125 times, and both are real products a real team might reasonably pick.

The discounts are large and underused. A cache hit on the Claude API costs 10 percent of the base input rate. The Batch API halves input and output together. Neither needs a different model or a better prompt, only a different request shape.

But the sticker price is one of two multiplicands. The other is how many tokens your text actually becomes, and that number is not the same for everybody.

The tokenizer charges Yoruba more than English for the same sentence

Tokenizer fertility is the count of tokens a tokenizer produces per word. It is not a constant. It is a property of what the tokenizer was trained on, which is overwhelmingly English and code.

Orevaoghene Ahia and colleagues put numbers on the consequence at EMNLP 2023 in a paper titled "Do All Languages Cost the Same?", measuring cost and utility across 22 languages on a commercial API and finding that speakers of many supported languages pay more while getting worse results. The finding has survived three years of model churn.

A 2026 preprint measuring 20 African languages across 11 tokenizers on the FLORES-200+ parallel corpus gives the current picture on o200k_base: English 1.22 tokens per word, Hausa 1.64, Igbo 1.73, Swahili 1.87, Yoruba 2.26, Tigrinya 8.27, Amharic 8.97 and Bambara written in N'Ko at 10.86. The median African premium is 1.88 times. Script dominates language family: Latin-script African languages average 1.76 times English, Ethiopic 7.08 and N'Ko 8.92.

Work presented at AfricaNLP 2026 on the AfriMMLU benchmark found something worse. Fertility predicts accuracy. Higher tokens per word means lower scores, consistently, across ten models and every subject tested. The languages that cost more also perform worse. Those are not independent problems.

If you want proof that token counts are a pricing lever rather than a fact of nature, read Anthropic's own footnote. Claude 4.7 and later ship a newer tokenizer that, per their pricing page, "produces approximately 30% more tokens for the same text". Same sentence. Same advertised price per token. Thirty percent more tokens. That is a price increase delivered through the tokenizer instead of the price sheet, disclosed in a note under a table.

Two taxes, one invoice

Put them together on a realistic workload. Take a support turn of 1,500 input tokens and 400 output tokens on gpt-5.6-terra at $2 and $12 per million. That is $0.0078 per turn in English.

Now price the product. ₦2,000 a month at ₦1,338 to the dollar is $1.49 of revenue. Allocate a quarter of it to inference and you have $0.37 per user per month, which buys about 48 turns.

Serve the same product in Yoruba. Fertility of 2.26 against English 1.22 means 1.85 times the tokens for identical meaning, so the turn costs $0.0144 and the budget buys 26 turns. In Amharic at 7.36 times, the turn costs $0.0574 and the budget buys six. Six turns a month.

Now move only the exchange rate. No new features, no worse prompts, no model swap. The naira returns to ₦1,600, which it has already done once inside the last three years, and the English user drops from 48 turns to 40.

The Amharic user was never viable at that price. The Yoruba user is viable at roughly half the headroom. And the founder who ships English-only is not making a product decision, they are making a currency decision they may not have noticed making.

The strongest objection: open weights make all of this temporary

The best argument against everything above is that per-token metering is a phase. Open weights sit close enough to frontier for most production work, tokenizers are improving quickly and self-hosting converts a variable dollar cost into an amortised one. The same 2026 study that produced the fertility numbers found Gemma 4's tokenizer averages 2.43 times English across 19 African languages against 3.40 for the older cl100k_base, a 28 percent reduction on identical text and 74.5 percent off on Amharic specifically. Picking a better tokenizer costs nothing. It is a config line.

That argument is right about the direction and wrong about the escape.

Self-hosting does not remove dollar exposure. It relocates it and gives it a worse shape. GPUs are priced in dollars, imported and cleared through customs. Power is the second input, and anyone who has run compute in Lagos knows the generator is not a footnote. You trade a variable cost that scales with revenue for a fixed cost that does not, which is precisely the wrong trade when volume is uncertain. A quiet month on an API is a small bill. A quiet month on a rack you have already paid for is the same bill.

The fertility penalty follows you across regardless. It belongs to the tokenizer, not the billing system. Self-host and it stops arriving as an invoice line and starts arriving as latency and as lost context. The same study measures it: at a 128k window, English fits about 105,000 words while Amharic fits about 14,000. You have not escaped the tax. You have changed which resource it eats.

Open weights change who you pay. They do not change what the tokenizer does to your language.

What actually moves the number

Measure fertility on your own corpus before you set a price, not after. Run real traffic through the candidate tokenizers and count. Published averages are a starting point rather than your number, and the gap between tokenizers on your specific text is frequently wider than the gap between models on your specific task.

Take the discounts. A cache read at 10 percent of input rate and a batch job at half price are not clever engineering, they are reading the pricing page all the way to the end.

Watch the clock, and notice whose clock it is. DeepSeek's documented peak hours run 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, with everything else at half price. Sitting in Shenzhen, that window reads as obviously local. It maps to 09:00 to 12:00 and 14:00 to 18:00 Beijing time, which is the Chinese working day with the lunch break cut out of the middle. Nobody drew that boundary with Lagos in mind. Convert it to West Africa Time anyway and a Lagos product gets an accident in its favour: peak falls at 02:00 to 05:00 and 07:00 to 11:00 WAT, so the entire Lagos afternoon and evening runs at the discounted rate. Late-day traffic is half price for reasons that have nothing to do with you.

The last one is not a tactic, it is something I keep noticing. DeepSeek publishes its price sheet twice: yuan per million tokens on the Chinese documentation, dollars per million on the English. The two track market FX closely, so it is not a discount. It is a currency of account. A developer in Shenzhen reasons about inference cost in the money they are paid in. There is no naira sheet. There is no cedi sheet. Every African builder does that conversion in their head, every time. They absorb the variance personally.

The decision nobody thinks of as a pricing decision

Fertility is fixed when the tokenizer is trained, by a corpus mix chosen months before anybody prices anything. That single choice sets, permanently, what a Yoruba sentence costs relative to an English one for the entire life of the model. It is a pricing decision made by people who do not experience it as one, in a room where nobody is thinking about naira.

Per-token prices are falling fast and will keep falling. The number of tokens your language costs is not falling at anything like the same rate, and on at least one frontier model it just went up 30 percent. Two curves, moving independently. Only one of them gets a launch post.

Tools referenced

vLLM, reviewed here: vLLM review.

Ollama, reviewed here: Ollama review.

llama.cpp, reviewed here: llama.cpp review.

LMDeploy, reviewed here: LMDeploy review.

Langfuse, reviewed here: Langfuse review.

DeepSeek, reviewed here: DeepSeek review.

Sources

Anthropic, Claude API pricing documentation: https://platform.claude.com/docs/en/about-claude/pricing

OpenAI, API pricing: https://developers.openai.com/api/docs/pricing

Google, Gemini API pricing: https://ai.google.dev/gemini-api/docs/pricing

DeepSeek, API models and pricing: https://api-docs.deepseek.com/quick_start/pricing

Ahia et al., "Do All Languages Cost the Same?", EMNLP 2023: https://aclanthology.org/2023.emnlp-main.614/

Lundin et al., "The Token Tax: Systematic Bias in Multilingual Tokenization", AfricaNLP 2026: https://aclanthology.org/2026.africanlp-main.10/

Somide, "The African Language Tax", arXiv preprint 2026: https://arxiv.org/abs/2606.24460

Monierate, CBN official USD/NGN rate: https://monierate.com/fx/official

Frequently Asked Questions

Why do African languages cost more to use with LLM APIs?

Because APIs bill per token, and tokenizers split African languages into far more tokens than English for the same meaning. On OpenAI's o200k_base tokenizer measured against the FLORES-200+ parallel corpus, English averages 1.22 tokens per word while Yoruba is 2.26, Amharic is 8.97 and Bambara written in N'Ko is 10.86. Since the price per token is identical for every language, an Amharic conversation costs roughly 7.4 times what the same conversation costs in English. The median premium across 20 African languages is about 1.88 times. Script matters more than language family: Latin-script African languages average 1.76 times English, while Ethiopic-script languages average 7.08 times.

How much do LLM API tokens cost in August 2026?

Published rates per million tokens as of 29 August 2026: Anthropic charges $5 input and $25 output for Claude Opus 5, $2 and $10 for Claude Sonnet 5 and $1 and $5 for Claude Haiku 4.5. OpenAI charges $4 and $20 for gpt-5.6-sol, $2 and $12 for gpt-5.6-terra and $0.05 and $0.40 for gpt-5-nano. Google charges $0.75 and $3.75 for Gemini 3.7 Flash through 31 December 2026, doubling afterwards. DeepSeek charges $0.22 input and $0.66 output for v4-flash off-peak, doubling during peak hours. Output token prices therefore span roughly 125 times between the cheapest and most expensive tiers, before any caching or batch discounts.

Does self-hosting an open-weight model remove currency risk for African startups?

No, it relocates the risk. GPUs, hosting and power are still priced in dollars or bought as imported hardware, and self-hosting swaps a variable cost that scales with revenue for a fixed cost that does not, which is the worse position when volume is uncertain. Self-hosting also does not fix tokenizer fertility, because fertility is a property of the tokenizer rather than the billing system. The extra tokens stop appearing as invoice lines and start appearing as latency and lost context window instead. At a 128k context window, English fits roughly 105,000 words while Amharic fits about 14,000. Choosing a better-fitting tokenizer helps more than changing who hosts the model.