Skip to content
💡 Innovation Ecosystem

What Happens When AI Gets Cheap, Seen From Shenzhen and Lagos

Francis Okafor Francis Okafor
10 min read
AI economics LLM pricing Manufacturing African markets Shenzhen Domain expertise Low-resource languages
What Happens When AI Gets Cheap, Seen From Shenzhen and Lagos
On this page
  1. The floor fell out and the ceiling went up at the same time
  2. What no model knows about mould number seven
  3. The price in Alaba is not the price in the training data
  4. Yoruba is not low-resource. It is unrecorded.
  5. The strongest argument against this: most expertise was retrieval
  6. The premium is visible on the price sheet
  7. The line moves, and writing about it is how you move it
  8. Tools referenced
  9. Sources

OpenAI shipped GPT-4 on 14 March 2023 at $30 per million input tokens and $60 per million output tokens. I checked the same company's price sheet this morning. gpt-5-nano: $0.05 in, $0.40 out. DeepSeek charges $0.007 per million tokens for a cached input read on v4-flash during off-peak hours. So if you want to know what happens when AI gets cheap, stop forecasting and read an invoice. It already happened.

Input pricing fell 600x in forty-one months. Output fell 150x. And the model at $0.05 is better at almost everything than the model at $30. That is the part most people still have not absorbed.

Here is what I think that did to the market for knowing things. The generic got commoditised. The specific got a premium. And the specific was never concentrated in San Francisco, Beijing or London, because specific knowledge is not a talent pool. It is a location.

The floor fell out and the ceiling went up at the same time

Two measurements, from the same dataset, both true.

Hans Gundlach, Jayson Lynch, Matthias Mertens and Neil Thompson at MIT FutureTech assembled the largest set of historical AI pricing data published so far. Their finding, in a paper first posted 28 November 2025 and revised 23 March 2026: the price for a given level of benchmark performance is falling roughly 5x to 10x per year across knowledge, reasoning, maths and software engineering tasks. Strip out competitive pricing and hardware improvements, and the underlying algorithmic efficiency gain sits near 3x a year.

Epoch AI measures the same effect from another angle. The price of matching GPT-4 on GPQA Diamond, the PhD-level science question set, fell about 40x per year. Across other capability milestones the annual decline ranges from 9x to 900x. Wide spread, one direction.

Now the second measurement, from that same MIT paper. The price of operating the frontier itself is rising 3x to 18x per year. Bigger models, longer reasoning traces, far more tokens burned per answer.

So getting a fixed job done keeps collapsing in price while standing at the front keeps getting more expensive. The current sheet shows both at once. Claude Haiku 4.5 sits at $1 in and $5 out. Claude Sonnet 5 is $2 and $10, an introductory rate Anthropic has confirmed becomes the standard price instead of rising to $3 and $15 on 1 September 2026. Claude Opus 5 is $5 and $25. gpt-5.6-sol lists at $4 and $20. DeepSeek v4-pro off-peak is $0.66 and $1.98. Between the cheapest usable tier and the dearest one, two orders of magnitude, same month, same market.

A detail almost nobody prices in: Claude 4.7 and later use a tokenizer that produces roughly 30% more tokens for the same text. The headline per-token number went down. The bill did not go down by as much.

At Nigeria's official window on 28 August 2026, around 1,341 naira to the dollar, a million input tokens on gpt-5-nano works out to about 67 naira. A million tokens is roughly 750,000 words. Whatever that is, it is not a moat.

The three items on the left are near-free to reason about because every model has read the documents behind them. The three on the right were never recorded, so no model has read them, and that is where the premium moved.
The three items on the left are near-free to reason about because every model has read the documents behind them. The three on the right were never recorded, so no model has read them, and that is where the premium moved.
The line does not run between shallow and deep. It runs between recorded and unrecorded.

What no model knows about mould number seven

Take an injection moulding tool that has been in production for four years in a plant in Bao'an. The cavity has worn. The gate is slightly eroded. Resin now arrives from a second supplier and that lot flows differently at the same barrel temperature. The operator bumps the barrel up three degrees after the morning break and has done so for two years without telling anyone, because that is what makes the part fill.

Ask a frontier model to write a statistical process control routine for that line and you get clean, correct code in about nine seconds for a fraction of a cent. Ask it what is actually wrong with the parts and it has nothing, because the input it needs was never written anywhere.

The most useful thirty seconds I have spent in a factory was watching a line leader put her flat palm on a jig before first shift. She was checking whether it had cooled overnight. No standard operating procedure told her to do that. Nothing in the MES logged it. She had worked out that a warm jig gives you a first article that passes inspection and then walks out of tolerance by the second hour, and that if you sign off on a warm jig you scrap parts until lunch.

That is a causal model of a physical system, held by one person, never recorded. There are hundreds of thousands of them inside Guangdong's factories. Every one is worth more now than in 2023, precisely because the thing it feeds became free.

Vision inspection makes the point brutally. Ultralytics YOLO26 weights cost nothing. Labelling the specific defects your line produces, on your parts, under your lighting, is the entire cost and the entire advantage. The model is a commodity. Your Label Studio project is not.

The price in Alaba is not the price in the training data

Nigeria has roughly 39.7 million micro, small and medium enterprises on SMEDAN and National Bureau of Statistics figures, of which close to nine in ten operate informally. The informal sector employs well over 70% of the labour force. Almost none of that transacts through anything a crawler can read. There is no API for Alaba International Market.

Ask a model what a 2 horsepower inverter sells for in Lagos and you get an average of scraped listings, which is a number nobody actually charges. The real number depends on whether you are buying one or forty, whether you are paying in naira or dollars and which rate the seller has decided to use today. On 28 August 2026 the official window quoted around 1,341 naira to the dollar while the parallel market sat near 1,405. That 5% gap is not a rounding error in a trade where a container clears on a 4% margin. It is the business.

I have watched a trader in Computer Village quote the same unit at three prices in one afternoon and be right every time, because he was pricing the buyer rather than the item. Knowing which of those three prices applies to you is the skill. It is absent from every dataset because it is absent from every document. It lives in his head and it changes by Thursday.

Yoruba is not low-resource. It is unrecorded.

AfroBench, from McGill NLP, evaluated 12 models across 64 African languages, 15 tasks and 22 datasets. GPT-4o scored more than 25 points better on English than on the African language average. Gemma 2 27B, more than 40. Knowledge-heavy and reasoning tasks showed the widest gaps. Nigeria alone has more than 500 living languages.

The safety side is now better documented. On 29 July 2026 the GSMA and Zindi published an African Trust and Safety LLM benchmark: 4,216 validated adversarial tests across eight languages, built by 307 contributors. Swahili takes 33.4% of the set, Hausa 21.6%, Yoruba 14.1%, Igbo 9.3%. A guardrail that holds in English need not hold in Hausa, because the guardrail training data barely exists in Hausa.

Mandarin is a high-resource language and the same failure appears in a different shape. 差不多 renders as "about the same". In a supplier message about a delivery date it does not mean about the same. It means the supplier has already decided the difference is acceptable and is informing you after the fact. Read it as a pleasantry and you have filed a schedule risk as a friendly exchange. Read it correctly and you make a phone call that afternoon. Machine translation gets the characters right and the message wrong, and the cost of that arrives six weeks later at the port.

Translation is now nearly free. Interpretation is not, and interpretation is where the money always was.

The strongest argument against this: most expertise was retrieval

Here is the objection I take seriously, and I think it is largely correct.

A great deal of what got called deep domain knowledge was memory plus lookup. Which clause of the standard applies. Which part number supersedes which. What the regulation says about that filing. What the failure mode was in 2019 and who fixed it. That is retrieval. Retrieval is exactly what now costs $0.05 per million tokens, and it does not sleep, take leave or resign.

Whole categories of paid work rested on being the person in the room who had read the document. First-pass contract review. Compliance summarisation. Tier one support. Boilerplate integration code. Market research decks. All repriced downward, permanently, and the people who did that work are not wrong to be angry. Telling them "deep domain knowledge is safe" is false comfort. Plenty of what gets sold as domain expertise is a wrapper around a PDF the model finished reading in 2024.

So the line does not run between shallow and deep. It runs between recorded and unrecorded.

If your knowledge sits in a manual, a forum thread, a standard, a filing or a stack of papers, it has been read and it is near-free. If it exists only as something a person observed with their hands on a machine, in a market, in a room, in a language nobody scraped, it has not been read. Seniority tells you nothing about which side you are on. I know engineers with twenty years who are squarely on the wrong side, and a 26-year-old line supervisor in Longhua who is squarely on the right one.

The test is uncomfortable and takes about a minute. Write down what you know that makes you valuable. Then ask whether a competent search would have found it. Whatever answers yes has already been repriced. What is left is your actual position.

The premium is visible on the price sheet

Look again at the spread between $0.05 and $25 per million tokens. That spread is a risk ladder, not a quality ladder. Nobody pays 100x for a summary. They pay 100x when being wrong is expensive and there is no reference to check against.

Which is why the frontier price climbs 3x to 18x a year while the price of a fixed benchmark score falls 5x to 10x. Buyers are sorting themselves. High-volume, well-specified, checkable work drops to the cheap tier or leaves the API entirely for self-hosted open weights through vLLM or Ollama, where marginal cost is electricity and a GPU you already own. Ambiguous, consequential, unrepeatable work goes to the expensive tier, and increasingly it goes there carrying a large pile of your own context.

That pile is the asset. Not the weights. The weights are on Hugging Face.

So are your evals, or they should be. Public benchmarks tell you how a model performs on GPQA. They tell you nothing about how it performs on your defect taxonomy or six years of supplier correspondence. Building that eval set with something like Promptfoo is unglamorous, cheap and the only measurement that maps onto your money.

The line moves, and writing about it is how you move it

None of this is stable.

Recordedness is a state, not a property, and the state changes the moment somebody writes it down. The line leader's palm on the jig was unrecorded until I put it in a paragraph. Now it is a sentence on a public page. It will be crawled. Within a training cycle it becomes a plausible completion, and the next engineer who asks a model about first-article drift on a moulding line may well be told to check whether the jig has cooled.

I traded a sliver of an advantage for a point I wanted to make. Everyone writing about their edge is running the same arithmetic, mostly without noticing.

Which leaves something I cannot resolve. The people holding the deepest unrecorded knowledge of a factory, a market or a language have every reason not to publish it and no way to be paid for it while it stays private. The ones who do publish get read, hired and cited, and their advantage decays in direct proportion to how well they wrote. There is no equilibrium in that. There is only a rate of decay, and the price of a million tokens falling underneath it.

Tools referenced

DeepSeek, reviewed here: DeepSeek review.

Claude, reviewed here: Claude review.

ChatGPT, reviewed here: ChatGPT review.

vLLM, reviewed here: vLLM review.

Ollama, reviewed here: Ollama review.

Ultralytics YOLO26, reviewed here: Ultralytics YOLO26 review.

Sources

OpenAI API pricing (current model rates): https://developers.openai.com/api/docs/pricing

Anthropic Claude model pricing documentation: https://platform.claude.com/docs/en/about-claude/pricing

DeepSeek API models and pricing: https://api-docs.deepseek.com/quick_start/pricing

Epoch AI: LLM inference prices have fallen rapidly but unequally across tasks: https://epoch.ai/data-insights/llm-inference-price-trends

Gundlach, Lynch, Mertens & Thompson, The Price of Progress (MIT FutureTech, arXiv): https://arxiv.org/abs/2511.23455

AfroBench: How Good are Large Language Models on African Languages? (McGill NLP, arXiv): https://arxiv.org/abs/2311.07978

GSMA & Zindi: African Trust & Safety LLM Benchmark, 29 July 2026: https://www.gsma.com/newsroom/blog/african-trust-safety-llm-benchmark-stress-testing-ai-safety-across-africas-languages-and-contexts/

TechCrunch, OpenAI releases GPT-4 (launch pricing, 14 March 2023): https://techcrunch.com/2023/03/14/openai-releases-gpt-4-ai-that-it-claims-is-state-of-the-art/

Frequently Asked Questions

How much does a million tokens cost in 2026?

It depends entirely on tier. As of August 2026, OpenAI's gpt-5-nano costs $0.05 per million input tokens and $0.40 per million output tokens, while gpt-5 costs $1.25 and $10.00. Anthropic's Claude Haiku 4.5 is $1 and $5, Claude Sonnet 5 is $2 and $10 and Claude Opus 5 is $5 and $25. DeepSeek v4-flash off-peak charges $0.22 for cache-miss input, $0.007 for cache-hit input and $0.66 for output. For comparison, GPT-4 launched in March 2023 at $30 input and $60 output per million tokens.

Is domain expertise still valuable now that AI is cheap?

Some of it is and some of it is not, and the dividing line is whether the knowledge was ever written down. Expertise that consisted of recall and lookup, such as knowing which clause of a standard applies or which part number supersedes another, has been repriced sharply downward because models have already read those documents. Expertise that exists only as unrecorded observation, such as how a specific machine drifts out of tolerance, what a specific market actually charges or what a phrase means in context in a language with little training data, has not been read by any model and commands a higher premium than it did in 2023.

Why are frontier AI models getting more expensive while AI overall gets cheaper?

Both trends are real and they measure different things. Research from MIT FutureTech published in November 2025 and revised in March 2026 found that the price to achieve a given level of benchmark performance falls roughly 5x to 10x per year, while the price of operating frontier models themselves rises between 3x and 18x per year because models grow larger and reasoning tasks consume far more tokens per answer. In practice this means yesterday's capability keeps getting cheaper while today's best capability keeps getting dearer, producing a price sheet spanning two orders of magnitude in the same month.