Is AI a Bubble? Yes, and the Chips Will Still Be Running
Francis Okafor
On this page
- "Is AI a bubble" is a question about prices, not chips
- The depreciation argument and where it breaks
- Circular financing is the tell
- What the token meters actually say
- The strongest case against my position
- Which track survives a repricing
- Cheaper compute reads differently from Lagos
- Tools referenced
- Sources
Is AI a bubble? Almost certainly yes. I build these systems for a living from an office in Shenzhen and I have never owned a share of Nvidia, Oracle or CoreWeave, which is the only reason my answer is worth reading. I have no position to talk up and nothing to hedge.
Financial bubbles and technological usefulness are independent variables. They can point in opposite directions for a decade at a time. British railway shares fell more than 65 percent from their mid-1840s peak by 1850. About a third of the mileage authorised between 1844 and 1846 was never laid. The 6,220 miles that were laid still carry trains today. The bubble was real. So was the track.
So the interesting question is not whether asset prices are wrong. They are wrong. The questions are which specific assets survive a repricing, what the real depreciation schedule is for silicon running at full utilisation and who is holding the paper when the marks come down.
"Is AI a bubble" is a question about prices, not chips
Start with the capex, because that is what the whole argument rests on. In their July 2026 results Amazon, Alphabet, Meta and Microsoft guided to somewhere between 720 and 745 billion dollars of capital expenditure for calendar 2026. Amazon around 220 billion. Alphabet 195 to 205 billion. Meta 130 to 145 billion. Microsoft tracking near 175 to 190 billion depending on how you treat its accounting adjustments. The same four spent roughly 413 billion in 2025. That is not growth. That is a step function.
Widen the frame. Morgan Stanley and Moody's put the total AI data centre build-out at about three trillion dollars. JPMorgan goes above five trillion once you include the power supply. Bain's Global Technology Report, published 23 September 2025, calculated that the industry needs two trillion dollars of annual AI revenue by 2030 to fund the compute it is projecting and forecast a shortfall of roughly 800 billion.
Now the revenue side, where I have to correct a figure that still circulates. The claim that AI-attributable revenue sits between 50 and 150 billion dollars is stale. Anthropic's annualised run rate reached 65 billion dollars at the end of July 2026, up from about 9 billion at the end of 2025. OpenAI reported 40 billion in August 2026, doubled from 20 billion at the close of last year. That is 105 billion from two companies, before Google, Microsoft or Meta, none of which break out AI revenue cleanly.
It is still not enough. Run rate is an annualisation of one good month rather than a year of collected cash, and both companies burn heavily. Set 105 billion of model-layer run rate against 720 to 745 billion of capex in a single year and you get roughly one dollar of visible AI revenue for every seven dollars of annual spending. Bulls will say infrastructure always front-runs revenue. They are right. The argument is about how far and for how long.
The racks can run flat out for nine years while the equity that paid for them prints zero.
The depreciation argument and where it breaks
Michael Burry made the sharpest version of the bear case in November 2025. He argued that hyperscalers are extending the useful lives of AI hardware on their books, understating depreciation by about 176 billion dollars across 2026 to 2028 and inflating reported earnings as a result. He put Oracle's 2028 earnings overstatement near 27 percent and Meta's near 21 percent. His premise was that Nvidia's product cycle runs two to three years while the books assume five or six.
The premise is wrong, and I say that as someone sympathetic to his conclusion. Product cycle is not economic life. In February 2025 Amazon shortened the useful life of a subset of its servers from six years to five, citing the pace of AI development, which is the opposite direction from the accusation. CoreWeave has signed A100 contracts running into 2029. The A100 launched in May 2020. That is a nine-year revenue life on a part four generations behind the front. Nvidia T4s from 2018 still rent.
The reason is power, not silicon. When the binding constraint is a substation and an interconnect queue measured in years, an old accelerator in a rack you have already energised beats an empty rack by an infinite margin. Old GPUs get demoted to inference, to smaller models, to overnight batch. They do not get skipped.
That does not make the hyperscalers safe. It relocates the risk. The danger is not that the chips stop working. The danger is that the price per GPU-hour falls faster than the depreciation schedule assumes, so the asset runs at full utilisation and still fails to earn back its book value. That is a revenue problem wearing an accounting costume. Burry may turn out directionally right for a reason he did not argue.
Circular financing is the tell
The part that genuinely worries me is not the size of the spending. It is the shape of the flows.
Nvidia signed a letter of intent to invest up to 100 billion dollars in OpenAI, tied to ten gigawatts of deployment. AMD granted OpenAI warrants over roughly 160 million shares at a cent each. Oracle and OpenAI signed a 300 billion dollar compute agreement. Microsoft and Nvidia put around 15 billion into Anthropic, which committed roughly 30 billion to Microsoft cloud and Nvidia chips. Nvidia holds about 5 percent of CoreWeave, which buys Nvidia GPUs.
Then the debt, which is where the exposure actually sits. Alphabet, Amazon, Meta and Oracle borrowed about 93 billion dollars in the US investment-grade bond market during 2025, roughly 6 percent of all issuance that year. Oracle sold 18 billion in a single day that September. Meta arranged more than 27 billion of debt for one campus through a special purpose vehicle. AI-related companies tapped debt markets for at least 200 billion in 2025, an undercount because so many deals are private, and outstanding private credit loans to AI-related borrowers passed 200 billion. Morgan Stanley expects 250 to 300 billion of hyperscaler and joint-venture issuance in 2026 alone, plus around 170 billion of data centre project finance loans, a 57 percent increase on the prior year.
Revenue that returns to the vendor who financed it is not demand. Call it inventory financing with better optics. I have no view on whether any particular deal is improper. I have a very firm view that a dollar counted at four points around a circle is still one dollar and that the accounting will not tell you which point was real until the circle breaks somewhere.
What the token meters actually say
Here is where I part company with the loudest bears.
Epoch AI tracks what it costs to reach a fixed capability level over time. Depending on the benchmark, that cost falls between 9x and 900x per year. GPT-3-level performance on MMLU cost 60 dollars per million tokens in November 2021 and 7 cents by February 2025.
Volume went the other way. On Alphabet's 22 July 2026 earnings call Sundar Pichai said the company's model APIs were processing approximately 22 billion tokens per minute, up from 16 billion the previous quarter, with cloud backlog at 514 billion dollars. On OpenRouter, one of the few neutral multi-vendor samples of real inference demand, agentic token consumption rose from 0.51 trillion to 7.3 trillion tokens a week between February and August 2026. Human consumption grew 2.8x over the same window.
Steeply falling unit prices alongside steeply rising volume is what elasticity looks like. A demand mirage looks like the opposite: prices falling while volume stalls, because the buyers were never there in the first place. That is not the picture in front of me.
The honest caveat is that tokens are not dollars. Nearly 70 percent of agentic tokens on OpenRouter come from cached prompts billed at a fraction of the standard rate. Backlog is a promise. And OpenRouter skews toward open-weight models that are less token-efficient, which inflates the raw count relative to what the frontier labs bill.
The strongest case against my position
The best argument against everything above deserves stating at full strength rather than being set up to fall over.
The volume growth is real but purchased. The marginal buyer of inference in 2026 is disproportionately another AI company funded out of the same capital stack that funds the compute. Agents burn tokens in retry loops and redundant context, so token count measures compute consumed rather than value delivered, and a 14x rise in agentic tokens may be partly a measurement of inefficiency. Meanwhile the deflation cuts against the owner of the asset. If revenue per token falls faster than volume rises, revenue per dollar of installed capex declines even while utilisation looks perfect. You end up with millions of very busy assets that never return their cost of capital, financed by floating-rate private credit against take-or-pay contracts from counterparties who are themselves lossmaking. Bain's 800 billion dollar shortfall already assumes generous adoption.
I cannot dismiss the second half of that. Deflation outrunning volume is a genuine mechanism by which useful infrastructure becomes uneconomic infrastructure, and no amount of engineering enthusiasm makes it go away.
What I would say is that this is an argument about return on capital rather than about whether the demand exists. Those are different failure modes with different victims. Fake demand kills the technology. Bad returns kill the owners and hand the technology to everyone else at a discount. Every measurement I can independently check points at the second. Teams meter their own spend now with tooling like Langfuse and route around expensive models on price, which is behaviour you only see from someone paying attention to a bill they intend to keep paying. Two labs went from a combined 29 billion dollars of annualised revenue at the end of 2025 to roughly 105 billion by August 2026. You can argue about the quality of run-rate accounting. You cannot produce that curve without customers.
Which track survives a repricing
In 1846 Parliament passed 263 Acts authorising 9,500 miles of new railway. Around 6,220 miles from the 1844 to 1846 authorisations were actually built, against a modern UK network of roughly 11,000 miles. A third of what was approved never happened. The companies that survived were mostly not the ones that raised the money.
Sort the AI build-out the same way. Substations, transmission interconnects, land, cooling shells and water rights survive nearly anything, because their useful life runs in decades and their scarcity is physical rather than contractual. GPUs survive as long as power is the constraint, which is exactly what a nine-year A100 contract is telling you. Frontier model weights depreciate faster than any chip and are already close to commodity one tier down.
What does not survive is equity in a single-tenant special purpose vehicle financed against one counterparty's promise to pay. Nor does the second-tier cloud with one customer, one silicon generation and floating-rate debt. Those get carried out first and quietly, in a footnote, months before anyone writes the retrospective.
Cheaper compute reads differently from Lagos
In Huaqiangbei I have watched the asking price on last-generation accelerator boards drift down month after month while the queue for the newest parts stayed long. On a factory floor in Dongguan I watched a vision inspection line run a quantised model on hardware that would embarrass anyone's benchmark chart, catching solder bridges on a night shift at a rate the human inspectors could not hold past hour six. Nobody in that building had an opinion about Nvidia's forward multiple.
The gap between the financial story and the operational one is far wider outside the capital centres. Nigeria has around 86 MW of operational data centre capacity with more than 320 MW under construction or planned. Operators self-generate power at 28 to 33 US cents per kWh because the grid delivers 5,000 to 6,000 MW against 13,000 MW of installed capacity. Kasi Cloud's first 5.5 MW phase at Lekki came online in April 2026. Local GPU hours are advertised below a dollar.
For a team in Lagos or Nairobi paying dollar-denominated cloud bills out of a currency that keeps sliding against the dollar, a repricing that dumps H100 and H200 capacity onto the secondary market at distressed rates is not a disaster. It is the first moment the marginal cost of trying something falls to what a local company can pay out of revenue rather than out of a grant. Cheap open weights already do part of that work. DeepSeek's V4 tiers, Tongyi Qianwen, Kimi from Moonshot AI and MiniMax M3 undercut frontier output pricing by an order of magnitude and serve perfectly well on vLLM or LMDeploy on rented hardware, or on llama.cpp and Ollama when the model is small enough to sit next to the machine it is inspecting. Huawei claims its Ascend 950PR delivers roughly 2.87 times the inference throughput of an H20 at about a quarter of the cost. I take vendor claims from every flag with salt. Chinese operators are designing around it regardless.
The two failure modes are not mutually exclusive and nobody gets to pick which one arrives. The racks can run flat out for nine years while the equity that paid for them prints zero, because the people who own the compute and the people who own the paper stopped being the same people somewhere around the third special purpose vehicle. That separation was the entire point of the structure. It is going to be the entire problem.
Tools referenced
vLLM, reviewed here: vLLM review.
LMDeploy, reviewed here: LMDeploy review.
llama.cpp, reviewed here: llama.cpp review.
Ollama, reviewed here: Ollama review.
DeepSeek, reviewed here: DeepSeek review.
MiniMax M3, reviewed here: MiniMax M3 review.
Sources
NVIDIA Q2 FY2027 CFO Commentary (SEC Form 8-K, 26 August 2026): https://www.sec.gov/Archives/edgar/data/0001045810/000104581026000073/q2fy27cfocommentary.htm
Alphabet Q2 2026 earnings call, Sundar Pichai's remarks (22 July 2026): https://blog.google/company-news/inside-google/message-ceo/alphabet-earnings-q2-2026/
Epoch AI, LLM inference price trends: https://epoch.ai/data-insights/llm-inference-price-trends
Bain & Company, $2 trillion in new revenue needed to fund AI's scaling trend (23 September 2025): https://www.bain.com/about/media-center/press-releases/20252/$2-trillion-in-new-revenue-needed-to-fund-ais-scaling-trend---bain--companys-6th-annual-global-technology-report/
Michael Burry warns of $176 billion depreciation understatement (November 2025): https://finance.yahoo.com/news/michael-burry-warns-176-billion-173613512.html
Bloomberg via Insurance Journal, The $3 Trillion AI Data Center Build-Out Becomes All-Consuming for Debt Markets (3 February 2026): https://www.insurancejournal.com/news/international/2026/02/03/856623.htm
The Decoder, agentic token usage jumps 14x on OpenRouter (23 August 2026): https://the-decoder.com/ai-is-becoming-ais-biggest-customer-as-agentic-token-usage-jumps-14x-on-openrouter/
Techpoint Africa, AI infrastructure in Nigeria: closing the GPU gap (25 March 2026): https://techpoint.africa/guide/ai-infrastructure-in-nigeria-gpu-gap/
Frequently Asked Questions
Is AI a bubble in 2026?
By the financial definition, yes. Amazon, Alphabet, Meta and Microsoft guided in July 2026 to between 720 and 745 billion dollars of capital expenditure for the calendar year, against roughly 413 billion in 2025, while the two largest model companies together run at about 105 billion dollars annualised. Morgan Stanley and Moody's put the total build-out near three trillion dollars. That gap is a valuation problem. It is not evidence that the technology is unused: inference volume and inference prices have moved in opposite directions for three straight years, which is elasticity rather than a mirage. A bubble in the equity and the debt is entirely compatible with compute that keeps running afterwards.
How long do AI GPUs actually last before they are obsolete?
Far longer than the two to three year product cycle implies, because product cycle and economic life are different things. CoreWeave has signed contracts for Nvidia A100s, launched in May 2020, that run into 2029. Nvidia T4s from 2018 still generate rental revenue. Amazon moved a subset of its servers from a six-year to a five-year useful life in February 2025. Old accelerators get demoted to inference and smaller models rather than retired, because when power and interconnect capacity are the binding constraints an energised old rack beats an empty new one. The real risk is not chips failing but the rental price per GPU-hour falling faster than the depreciation schedule assumes.
What is circular financing in AI and why does it matter?
Circular financing is when a chip vendor funds its own customers, who then use that money to buy its chips. Nvidia signed a letter of intent to invest up to 100 billion dollars in OpenAI tied to ten gigawatts of deployment; OpenAI signed a 300 billion dollar compute agreement with Oracle; Nvidia holds about 5 percent of CoreWeave, which buys Nvidia GPUs; AMD granted OpenAI warrants over roughly 160 million shares at a cent each; and Microsoft and Nvidia invested around 15 billion in Anthropic, which committed about 30 billion back to Microsoft cloud and Nvidia chips. It matters because the same dollar can be booked as revenue at several points, which makes demand look larger and more independent than it is, and because much of the underlying build is funded by private credit and off-balance-sheet vehicles rather than by the operating companies themselves.