Skip to content
CHINA TECH FROM THE INSIDE intermediate

China-Available LLM APIs From Mainland Networks: The Real Map

Three gates decide what you can call from a mainland connection: the network, the account and the payment rail. Only one is a firewall, and it is the one that matters least.

September 2, 2026
9 min read
Francis Okafor
China-Available LLM APIs From Mainland Networks: The Real Map

The map of China-available LLM APIs from mainland connections is drawn almost entirely by people who have never used one. It shows a wall. American models on one side, Chinese models on the other, the Great Firewall in between. That picture is wrong in the way that costs money.

There are three gates, not one. Network reachability. Account creation. Payment. Only the first is a firewall question, and for the endpoints most people care about it is the gate that matters least.

I live in Shenzhen and hit all three every week. What follows is the structure as the providers themselves document it, checked on 1 September 2026. Where I tell you something is a policy, I quote the policy. Where I tell you something needs testing from your own network, I tell you how to test it rather than handing you a latency number you cannot reproduce.

The American APIs are closed by their own terms, not by the firewall

OpenAI's supported-countries page carries this sentence: "Accessing or offering access to our services outside of the countries and territories listed below may result in your account being blocked or suspended." Mainland China, Hong Kong and Macau are not on the list. OpenAI moved from a written policy to active blocking of traffic from unsupported regions in July 2024, and several Chinese labs ran migration promotions the same week.

Anthropic publishes an equivalent list covering both its commercial API and Claude.ai. Mainland China, Hong Kong and Macau are absent from both. Taiwan is on both. Google states that "The Gemini API and Google AI Studio are available in the following countries and territories" and mainland China does not appear.

The practical consequence is the part people skip. A tunnel does not solve this. It converts a network failure into a terms violation, and the penalty named in OpenAI's own text is the account. For a company the account holds the billing history, the rate limits, the fine-tunes and the organisation keys. Nobody sane risks that asset to avoid a migration. That is why serious teams here did not quietly proxy their way around it. They moved.

The same mainland connection reaches all three branches. What separates them is not the network but the account: the American branch is refused by provider policy, and the two Chinese branches differ by identity requirement
The same mainland connection reaches all three branches. What separates them is not the network but the account: the American branch is refused by provider policy, and the two Chinese branches differ by identity requirement
Peak hours are Beijing office hours. A Lagos team running batch work from late morning pays half what a Shenzhen team pays for identical tokens at the same instant.

China-available LLM APIs from mainland networks come in matched pairs

The structural fact that explains almost everything else is that nearly every Chinese lab now runs two API businesses. Two endpoints, two account systems, two price lists and two model catalogues.

Alibaba documents it most plainly. Model Studio serves Beijing (cn-beijing) behind dashscope.aliyuncs.com and Singapore (ap-southeast-1) behind dashscope-intl.aliyuncs.com, with both migrating to workspace-scoped hosts of the form {WorkspaceId}.cn-beijing.maas.aliyuncs.com and {WorkspaceId}.ap-southeast-1.maas.aliyuncs.com. The regions page states it without hedging: "Each region has its own access domain, API Key, and model list. These cannot be used across regions." Singapore serves the International deployment scope only, Beijing the Chinese mainland scope only. Fine-tuning is listed as available in Beijing and not in Singapore.

The pattern repeats across the field. Zhipu answers at open.bigmodel.cn/api/paas/v4 for the mainland and at api.z.ai/api/paas/v4 for everyone else, where the pricing page says "All prices are in USD." Moonshot serves api.moonshot.ai/v1 internationally and has already rebranded the console: platform.moonshot.ai returns a 301 to platform.kimi.ai. MiniMax splits api.minimax.io from api.minimaxi.com and exposes an Anthropic-shaped path at api.minimax.io/anthropic. ByteDance keeps Doubao inside Volcengine on ark.cn-beijing.volces.com/api/v3 and sells the same model family abroad through BytePlus ModelArk on ark.ap-southeast.bytepluses.com. SiliconFlow documents api.siliconflow.cn/v1 in one doc set and api.siliconflow.com/v1 in the other.

DeepSeek is the exception, and it shows how deliberate the rest of the field is being. One endpoint, api.deepseek.com, for everybody, an Anthropic-compatible path at api.deepseek.com/anthropic and a single price list quoted in US dollars.

The +86 number blocks the account, not the packets

The registration barrier is a legal artefact, not an engineering one. China's Cybersecurity Law imposes real-identity obligations on network service providers, and the Interim Measures for the Management of Generative Artificial Intelligence Services, in force since 15 August 2023, place generative AI providers inside that regime. So the mainland consoles want a mainland mobile number and then real-name verification against an identity document.

For a foreigner resident here that is friction rather than a wall. Carriers issue SIMs against a passport and a residence permit, and the cloud consoles accept passport-based individual verification; Alibaba's own account documentation instructs individuals outside the Chinese mainland to submit passport information. What is not negotiable is the account split. Alibaba states that "The account data of the two sites is completely isolated and cannot be shared" and that "an account cannot be changed from one to the other."

So the honest answer to how to use Chinese AI models without a Chinese phone number is that you do not defeat the phone requirement. You buy the other product. The international endpoints take an email address and an international card, and they are a different SKU with a different catalogue and different prices, not a clever workaround.

For a Nigerian or Kenyan company that matters more than it first appears, because there is no upgrade path. Starting on the international stack and later wanting mainland latency or mainland pricing means registering from zero on the China site, with a Chinese business licence behind it, which in practice means a WFOE or a local partner.

The same model does not cost the same money on the two sides

Alibaba prices both regions in USD on its billing page, which makes the comparison unusually clean. The flagship qwen3.8-max is listed at $2 per million input tokens and $6 per million output in Singapore, against $1.65 and $4.951 for the same model in Beijing. Roughly a fifth cheaper inside the mainland.

The compensation runs the other way at the door. The same page: "The following models offer a free quota only in Singapore. No free quota is available in other regions." One million tokens, valid ninety days. The side that is easier to join is the more expensive side to run. That is a fine trade while you are evaluating and a bad one at volume.

DeepSeek prices in USD for everyone and discounts by the clock instead. Its pricing page puts deepseek-v4-pro at $1.32 per million input tokens on a cache miss and $3.96 per million output at peak, halving to $0.66 and $1.98 off-peak, with peak defined as "01:00 - 04:00 and 06:00 - 10:00 UTC, Monday through Friday". Cache hits fall by another order of magnitude, to $0.022 per million off-peak.

Convert those hours and something useful drops out. Peak is 09:00 to 12:00 and 14:00 to 18:00 in Shenzhen, which is exactly the Chinese working day. In Lagos the same window is 02:00 to 05:00 and 07:00 to 11:00. "A Lagos team that runs its batch work from late morning onward spends most of its day in the off-peak half of the price list, while a Shenzhen team working its own business hours spends most of its day in the peak half. The rate at any single instant is the same for both. What differs is which instants each team is awake for." "The discount was designed to flatten China's own load curve. West Africa lands mostly in the trough by accident, though not entirely: the second peak block runs to 11:00 Lagos time, so the first couple of hours of a normal working day there are still billed at the peak rate. Schedule the batch work after 11:00 and the rest of the day is half price."

Payment stops more foreign teams than the firewall does

Alibaba states the money split flatly: the China site settles in Chinese Yuan, the International site in US dollars with other major currencies supported. The mainland consoles expect Alipay, WeChat Pay or a mainland bank card. The international consoles expect an ordinary international card.

Alipay will now bind foreign Visa and Mastercard, which is how most foreigners here pay for everything, but it carries per-transaction and annual ceilings. Check the current limits inside the app before you make it a production dependency. A top-up that fails at a ceiling on the twentieth of the month is an expensive way to discover where the ceiling was.

Then there is paperwork, which decides more procurement arguments than latency ever will. A mainland purchase produces a fapiao, which your Lagos or Nairobi accountant has never seen and cannot file. An international purchase produces a normal invoice from a Singapore entity in dollars. For an African company with no China entity that single fact usually settles the question before any technical case is heard.

The strongest argument against doing any of this

Data is the real objection and it deserves a straight answer. Anything you send to a mainland endpoint is processed under Chinese law, and the provider carries content obligations under the Interim Measures that attach to what your users type and to what the model says back. Plenty of buyers, public sector buyers especially, will refuse that outright. They are not being irrational.

The obvious reply is that the international endpoints exist precisely for this. That reply is weaker than it looks. Alibaba's documentation describes Singapore as a region and a service deployment scope, which is a statement about where inference runs and where data sits. It is not a statement about corporate structure or about which laws reach the parent company. If your compliance position depends on that distinction, put it in a contract rather than inferring it from a docs page.

The clean exit is weights. DeepSeek, Qwen, GLM, Kimi and MiniMax all publish open models, and self-hosting collapses the endpoint question, the account question and the payment question into one procurement decision about GPUs. The wrinkle from this side of the border is supply: huggingface.co is unreachable from a mainland connection, and GreatFire's tracker still records it as blocked with tests logged into August 2026. Inside China you pull from ModelScope or a mirror. Outside China you pull from Hugging Face. Same weights, two supply chains, and only one of them is documented in the average README.

What I actually do, and what I tell African teams to do

Run two accounts and never pretend one is the other. A mainland account for anything latency-sensitive serving users inside China, an international account for everything else. Keys never cross, because the providers have already made sure they cannot.

Do not proxy an American API into production from here. The terms name the account as the penalty and the account is the asset you cannot re-buy.

Put base URLs in configuration, never in code. Inside the window this piece covers, one console domain had already redirected to a new brand and one provider was mid-migration to workspace-scoped hosts. Endpoints move faster than SDK releases do.

Keep an OpenAI-shaped client and change two strings. Almost every provider here ships an OpenAI-compatible surface, and DeepSeek and MiniMax now ship Anthropic-compatible paths as well, so an agent harness written against either SDK will run against Chinese models with a base URL and a key swap.

Test from the network you will deploy on, not from your laptop. One timed request per candidate endpoint, issued from the actual production egress and logged with a date, is worth more than any table on the internet, this one included. An office fibre line in Shenzhen, a China Mobile APN and a mainland cloud VPC are three different questions with three different answers.

Budget in the currency the console bills in, and set a spend alert well below whatever payment ceiling you are relying on.

This map has a shelf life measured in weeks

Everything above was checked against provider documentation on 1 September 2026. The model names alone will date it. DeepSeek's price list currently names deepseek-v4-flash and deepseek-v4-pro, Alibaba's catalogue lists qwen3.8-max, Kimi's docs list K3. None of those strings will be correct in a year and some will be wrong by Christmas.

The structure outlasts the strings. Two stacks per lab, separate accounts, separate money, a phone number at one door and a passport at the other. That arrangement has survived several model generations and there is no commercial reason for it to stop, because it lets the same company sell under two legal regimes without pretending they are one.

The wall everyone draws on the map is not where the border is. The border is a phone number, a business licence and a payment rail, and none of the three shows up in a traceroute.

Tools referenced

ModelScope, reviewed here: ModelScope review.

Sources

OpenAI, Supported countries and territories (API documentation): https://developers.openai.com/api/docs/supported-countries

Anthropic, Supported countries and regions: https://www.anthropic.com/supported-countries

Google, Gemini API available regions: https://ai.google.dev/gemini-api/docs/available-regions

Alibaba Cloud Model Studio, Select region, service deployment scope, and access domain: https://www.alibabacloud.com/help/en/model-studio/regions/

Alibaba Cloud Model Studio, Billing and model pricing: https://www.alibabacloud.com/help/en/model-studio/billing-for-model-studio

Alibaba Cloud, China site (aliyun.com) versus International site (alibabacloud.com): https://www.alibabacloud.com/help/en/account/aliyun-vs-alibaba-cloud

DeepSeek API documentation, Models and pricing: https://api-docs.deepseek.com/quick_start/pricing

Z.ai API documentation, HTTP interface introduction (base URL): https://docs.z.ai/guides/develop/http/introduction

Frequently Asked Questions

Can I use Chinese AI models without a Chinese phone number?

Yes, by using the international endpoints rather than the mainland ones. DeepSeek runs a single platform at api.deepseek.com that accepts email registration. Alibaba's Qwen models are reachable through the Singapore region at dashscope-intl.aliyuncs.com on an alibabacloud.com account, Zhipu's GLM models through api.z.ai, Moonshot's Kimi models through api.moonshot.ai and ByteDance's models through BytePlus ModelArk. These accounts take an email address and an international card. They are a separate product from the mainland platforms, with different model lists and different prices, not a way around the +86 requirement.

Does the OpenAI API work from mainland China?

No, and the blocker is OpenAI's own policy rather than only the firewall. OpenAI's supported-countries page states that accessing its services outside the listed countries and territories may result in an account being blocked or suspended, and mainland China, Hong Kong and Macau are not listed. OpenAI began actively blocking traffic from unsupported regions in July 2024. Anthropic's supported-countries list also excludes mainland China, Hong Kong and Macau, and Google does not list mainland China among Gemini API regions. Routing through a tunnel converts a network problem into a terms problem with the account as the stake.

What is the difference between dashscope.aliyuncs.com and dashscope-intl.aliyuncs.com?

They are two different regions of Alibaba Cloud Model Studio with separate everything. dashscope.aliyuncs.com is the China (Beijing) region, serving the Chinese mainland deployment scope and billed in Chinese Yuan on an aliyun.com account. dashscope-intl.aliyuncs.com is the Singapore region, serving the International scope and billed in US dollars on an alibabacloud.com account. Alibaba's regions documentation states that each region has its own access domain, API key and model list, and that these cannot be used across regions. Both are migrating to workspace-scoped hosts under maas.aliyuncs.com.

Are mainland and international API keys interchangeable?

No. For Alibaba Cloud Model Studio the documentation is explicit that an API key created in one region is rejected by another region's base URL, and that China site and International site accounts are completely isolated and cannot be converted into one another. The same separation applies in practice across Zhipu (open.bigmodel.cn versus api.z.ai), MiniMax (api.minimaxi.com versus api.minimax.io) and ByteDance (Volcengine Ark versus BytePlus ModelArk). Plan for two accounts and two billing relationships from the start, because there is no migration path between them.

Is it cheaper to call Chinese models from inside the mainland?

For Alibaba's Qwen models, yes, by roughly a fifth. Alibaba's own billing page prices qwen3.8-max at $2 per million input tokens and $6 per million output in the Singapore region, against $1.65 and $4.951 for the same model in Beijing. The international side compensates at the entrance: Alibaba states that a free quota of one million tokens, valid ninety days, is offered only in Singapore. DeepSeek instead charges everyone the same and discounts by time of day, halving its rates outside 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays.

How does an African company pay for a Chinese AI API?

Through the international endpoints, with an ordinary company card, receiving a US dollar invoice from a Singapore entity that a normal finance department can process. The mainland platforms bill in Chinese Yuan and expect Alipay, WeChat Pay or a mainland bank card, and they issue a fapiao rather than an invoice. Alipay will bind foreign Visa and Mastercard for individuals, but it carries per-transaction and annual ceilings that make it a poor foundation for a production API bill. Buying on the mainland side as a company also requires a Chinese business licence, meaning a local entity or partner.

Read next

China's University Major Cuts Are AI Policy, and Nigeria Should Read the Fine Print

The latest analysis essay.

Keep reading

Working on something in this space?

If this analysis is close to a problem you're thinking about, say so. I read every message personally.

Start a conversation