Comparison
Self-hosted inference servers vs hosted model APIs in 2026
Groq serves gpt-oss-120b at the same rate that puts the self-hosting break-even at 72 percent utilisation. So the real choice is not price. It is which jurisdiction your tokens are allowed to land in.
Last reviewed · 4 tools · 8 criteria
Verdict
Call a hosted API unless you have a residency rule, an export-control problem or a flat batch load, because the utilisation arithmetic almost never favours the other choice. When you do self-host, run vLLM: wider model coverage, a larger contributor base and in-tree support for NVIDIA, AMD and Intel. Run LMDeploy instead when your accelerators are Huawei Ascend, Cambricon or MetaX, where its single build system and named Atlas device support are a shorter path than assembling plugin backends. On the hosted side the two are not substitutes and price is not the deciding variable: Groq stores customer data in US buckets, DeepSeek stores it in the People's Republic of China, and one of those sentences usually disqualifies one of them before you reach the price card.
Side by side
| Criterion | vLLM | LMDeploy | Groq | DeepSeek |
|---|---|---|---|---|
| Licence on the thing you run | Apache 2.0 | Apache 2.0 | Proprietary service, no self-host path | Proprietary API; weights published separately under MIT |
| Hardware it runs on | NVIDIA, AMD and Intel GPUs plus x86/ARM/PowerPC CPUs in tree; plugin backends for Google TPU, Intel Gaudi, IBM Spyre, Huawei Ascend, Rebellions NPU, Apple Silicon and MetaX | NVIDIA CUDA plus build targets for Ascend, Cambricon, MACA and ROCm from one tree; Ascend support names Atlas 800T A3, Atlas 800T A2 and Atlas 300I Duo | Groq LPU only, not sold to customers | Not applicable, hosted only |
| Published price per million tokens | None, you pay for GPU hours | None, you pay for GPU hours | gpt-oss-120b $0.15 in / $0.60 out; gpt-oss-20b $0.075 in / $0.30 out | deepseek-v4-flash $0.44 in cache miss / $1.32 out at peak, $0.22 / $0.66 off-peak; v4-pro $1.32 / $3.96 peak, $0.66 / $1.98 off-peak |
| Discount structure | Not applicable | Not applicable | Batch and prompt caching discounts documented in Console | Off-peak at 50 percent; peak is 01:00-04:00 and 06:00-10:00 UTC Monday to Friday, which is 09:00-12:00 and 14:00-18:00 Beijing |
| Where the data physically sits | Wherever your GPUs are | Wherever your GPUs are | "All customer data is retained in Google Cloud Platform (GCP) buckets located in the United States"; inference logs up to 30 days, Zero Data Retention toggle available | "we directly collect, process and store your Personal Data in People's Republic of China" |
| Reach from mainland China and from African markets | Runs on your own hardware in either | Runs on your own hardware in either; the Ascend path matters where NVIDIA is unavailable | No site in Africa; nearest infrastructure is Dammam, Saudi Arabia and Helsinki. US export control and sanctions clauses in the services agreement | Reachable from both; servers in China, no African region |
| Release cadence you have to track | v0.27.0, v0.27.1 and v0.28.0 all shipped inside one month | v0.15.0 on 31 July, v0.16.0 on 19 August, v0.17.0 on 1 September | Vendor's problem | Vendor's problem |
| Throughput or availability commitment | None, it is yours | None, it is yours | Services agreement defers to service level commitments in separate Console terms rather than stating one | Concurrency caps of 500 (v4-pro) and 2,500 (v4-flash), HTTP 429 above; capacity expansion by request; no SLA in the rate-limit docs |
Which one to pick
-
vLLM
You have decided to self-host, your fleet is NVIDIA, AMD or Intel, and you want the widest model coverage and the largest contributor base. It is the default self-hosted server and the burden of proof is on anything else.
-
LMDeploy
Your accelerators are Huawei Ascend, Cambricon or MetaX, whether by procurement policy or because export control leaves you no alternative. One build system covers all of them and the Ascend documentation names supported Atlas devices outright, which is a shorter path than assembling plugin backends.
-
Groq
You are latency-bound on open-weight models, US data storage is acceptable or you will switch on Zero Data Retention, and you want Llama licence obligations to be someone else's compliance surface. Not an option if you need an endpoint reachable from mainland China.
-
DeepSeek
You are operating in or selling into mainland China, or you are outside it and your traffic concentrates in the off-peak band, where deepseek-v4-flash at $0.22 input and $0.66 output undercuts almost everything. Rule it out immediately if data may not be stored in the People's Republic of China.
Two rented H100s serving gpt-oss-120b cost $4,803 a month. Groq lists the same weights at $0.15 per million input tokens and $0.60 per million output, which is the identical rate I used when I worked the break-even out in full. It landed at 72 percent sustained utilisation, and almost nobody's traffic holds that. I am not redoing the arithmetic here. It is in The break-even on self-hosted open weights is 72 percent utilisation, and nothing since has moved it.
What that piece did not settle is which software you run on the two H100s, or which endpoint you call when you decide not to buy them. Those are separate questions with separate inputs. Cost decides whether to self-host. Geography and law decide who you call instead, and from Shenzhen or Lagos that geography is not the one the pricing pages assume.
Four tools from the review set frame it. vLLM and LMDeploy sit on the side where you own the hardware. Groq and DeepSeek sit on the side where you own an API key. Every figure below comes from vendor documentation, a privacy policy or a licence file, checked on 1 September 2026.
The two hosted APIs are not substitutes
Groq runs open-weight models on its own LPU silicon. The production catalogue is Meta's Llama 3.1 8B and Llama 3.3 70B, OpenAI's gpt-oss-120b and gpt-oss-20b, and Whisper Large V3 for speech. gpt-oss-20b is $0.075 input and $0.30 output per million. There is no self-hosted Groq. You cannot buy an LPU and put it in a rack in Lagos, so the hardware question does not exist and the residency question is answered for you.
DeepSeek's price card is the more interesting one, and the part that English-language comparisons consistently miss is the clock. Peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday through Friday. Everything else bills at half. Convert to Beijing time and those two windows are 09:00 to 12:00 and 14:00 to 18:00, which is the Chinese working day with the lunch break cut out of the middle. The discount is not a generic night rate. It is a load-shaping mechanism aimed at domestic office hours.
That has a direct consequence for anyone billing from West Africa. Lagos is UTC+1, so the peak windows land at 02:00 to 05:00 and 07:00 to 11:00 local. A Nigerian team whose traffic concentrates after lunch spends its entire working afternoon in the off-peak band, paying $0.22 per million input on a cache miss and $0.66 output for deepseek-v4-flash rather than $0.44 and $1.32. Nobody designed that discount for Lagos. It arrives there anyway, and it is worth more than most of the optimisations people spend a sprint on.
Neither vendor commits to much. DeepSeek publishes concurrency caps of 500 for v4-pro and 2,500 for v4-flash, returns HTTP 429 above them and invites a capacity expansion request, with no SLA in the rate-limit documentation. Groq's services agreement points at service level commitments living in separate Console terms rather than stating one.
vLLM is the default, LMDeploy is the one that boots on Ascend
Both are Apache 2.0, so the licence is not the differentiator. Coverage is. vLLM's README lists NVIDIA, AMD and Intel GPUs plus x86, ARM and PowerPC CPUs, with plugin backends for Google TPUs, Intel Gaudi, IBM Spyre, Huawei Ascend, Rebellions NPU, Apple Silicon and MetaX. Ascend support lives in vllm-ascend, a separate community-maintained repository with its own release train.
LMDeploy comes from Shanghai AI Laboratory's InternLM group and takes the opposite approach. Two engines, TurboMind in C++ and CUDA for throughput and a pure-Python PyTorch engine for anything TurboMind has not absorbed yet. Accelerator support is a build-time switch on LMDEPLOY_TARGET_DEVICE covering CUDA, Ascend, Cambricon, MACA and ROCm from one tree. The Ascend documentation names the devices outright: Atlas 800T A3, Atlas 800T A2 and Atlas 300I Duo, with prebuilt Docker images including a Kunpeng CPU variant.
Concede the obvious. If your fleet is NVIDIA, vLLM wins on merit and there is no serious argument for LMDeploy. Larger contributor base, faster model enablement, more eyes on regressions. The argument for LMDeploy appears at exactly the point where the accelerator is domestic, and in Shenzhen that point arrives more often than the English-language discussion suggests, because an H100 is not on the menu at any price and an Atlas card is.
Operational burden is the same shape for both, and it is a release-notes problem. vLLM shipped v0.27.0, v0.27.1 and v0.28.0 inside a single month. LMDeploy shipped v0.15.0 on 31 July, v0.16.0 on 19 August and v0.17.0 on 1 September. Both projects assume you follow upstream weekly. That is the quarter of an engineer that costs roughly what the GPUs cost, and it does not go away because you picked the calmer-looking project.
Latency and reach, from Shenzhen and from Lagos
Groq's footprint is US colocation with Equinix and DataBank, Bell Canada, Helsinki through Equinix and Dammam in Saudi Arabia with HUMAIN. The Dammam site is marketed as the largest AI compute hub across Europe, the Middle East and Africa, and it is worth being precise about what that means: it is the nearest thing to an African region and it is not in Africa. There is no Groq site on the continent.
The frontier vendors draw a different map. OpenAI's supported countries list includes Nigeria and excludes both mainland China and Hong Kong. Anthropic's list includes Nigeria, South Africa and Kenya, and excludes China and Hong Kong. Google has no mainland presence at all, though Google Cloud does operate africa-south1 in Johannesburg and lists it among locations for machine learning services, which makes it the only frontier-vendor region physically on the continent.
So the two markets fail in opposite directions. From Shenzhen the frontier APIs are closed as a matter of vendor policy, and the practical set is domestic endpoints or your own hardware. From Lagos the frontier APIs are legally open and physically distant, with the nearest region in Frankfurt or Northern Virginia and roughly a tenth of a second of round trip before the model has read the prompt. Invisible in a chat completion. Seconds per turn in an agent making a dozen sequential tool calls, which is why the latency argument for self-hosting is stronger in Nigeria than the cost argument ever was.
Residency comes down to one sentence in each policy
Groq's data documentation states that all customer data is retained in Google Cloud Platform buckets located in the United States. Inference logs are kept up to 30 days by default for reliability and abuse monitoring, and customers can switch on Zero Data Retention in Data Controls to stop that.
DeepSeek's privacy policy is equally direct in the other direction. It says the company directly collects, processes and stores personal data in the People's Republic of China, and repeats the sentence in the EEA supplement. There is no regional endpoint that changes this.
For a Nigerian bank under central bank data rules, a hospital group or a ministry handling court filings, those two sentences settle the matter before anyone opens a price card. One of them is disqualifying and the other may be too. Self-hosting is the only configuration where you answer the residency question yourself, which is the honest reason most on-premises inference clusters exist. It is not that the arithmetic favoured them.
The licence question is about the weights, not the server
vLLM and LMDeploy are both Apache 2.0 and neither constrains what you serve on them. The obligations travel with the checkpoint, and they vary more than the phrase open weights implies.
Qwen3-235B-A22B is Apache 2.0 with nothing attached. DeepSeek publishes its weights under MIT. Meta's Llama 4 Community License is a different instrument: you must display "Built with Llama" prominently on a related website or product documentation, any model you train on Llama outputs must begin with "Llama" in its name, and above 700 million monthly active users you need a separate licence that Meta grants at its sole discretion. MiniMax's community licence sets a revenue threshold and attribution terms of its own.
Calling Groq removes all of that from your compliance surface. You are buying inference under Groq's terms, and whatever the model licence requires of the operator is Groq's problem rather than yours. For Llama specifically that is an underrated argument for the hosted path, and it is one that never appears in a cost comparison because it does not have a dollar figure.
Where neither answer is right
Below a handful of concurrent users, both serving engines are the wrong tool. A single developer running a quantised model on a laptop, an air-gapped evaluation, a demo that has to work on a plane: Ollama and llama.cpp are both MIT, both start in one command, and neither asks you to reason about continuous batching or paged attention. Standing up vLLM for one user is a configuration burden with no throughput payoff.
The other case where the comparison collapses is the one I keep meeting in Huaqiangbei. A small firm buying a memory-modified 4090 for private deployment is not choosing self-hosting over an API. The API is closed to them by export control and vendor policy, and the modified consumer card is what is on the shelf. Huawei's Ascend line is the institutional version of the same trade, bought for availability rather than performance per dollar, and it is precisely why LMDeploy's single build system matters to a set of buyers who never appear in an English-language comparison table.
Everyone else should call the API and get the pager rotation back.
Questions
Is running vLLM cheaper than calling a hosted API?
Only above a utilisation threshold most organisations never reach. Groq lists gpt-oss-120b at $0.15 per million input tokens and $0.60 output, the same rate that puts the break-even for two rented H100s at 72 percent sustained utilisation. A business-hours workload sized for peak typically runs far below that. The full arithmetic is in the break-even piece, and the software you choose does not change it: vLLM and LMDeploy are both free.
Should I use vLLM or LMDeploy?
vLLM unless your accelerators are Chinese. It has wider model coverage, a larger contributor base and in-tree support for NVIDIA, AMD and Intel GPUs plus x86, ARM and PowerPC CPUs. LMDeploy earns the switch when you are on Huawei Ascend, Cambricon or MetaX, because a single build-time switch on LMDEPLOY_TARGET_DEVICE covers all of them and its Ascend documentation names Atlas 800T A3, Atlas 800T A2 and Atlas 300I Duo with prebuilt images.
Can I call Groq, OpenAI or Anthropic from mainland China?
Not on the vendors' own terms. OpenAI's supported countries and territories list excludes mainland China and Hong Kong, and so does Anthropic's. Google has no mainland presence. Groq's services agreement carries US export control and sanctions clauses and the company operates no site in the region. Inside China the practical options are domestic endpoints such as DeepSeek, or your own hardware.
Where is my data stored with Groq versus DeepSeek?
Groq's documentation states that all customer data is retained in Google Cloud Platform buckets located in the United States, with inference logs kept up to 30 days by default and a Zero Data Retention toggle available. DeepSeek's privacy policy states that it directly collects, processes and stores personal data in the People's Republic of China. For regulated buyers one of those two sentences is usually disqualifying before price enters the conversation.
Does the open-weight licence affect which server I run?
No. vLLM and LMDeploy are both Apache 2.0 and neither constrains what you serve. The obligations travel with the checkpoint. Qwen3-235B-A22B is Apache 2.0 and DeepSeek's weights are MIT, but Meta's Llama 4 Community License requires a prominent "Built with Llama" notice, requires derived model names to begin with "Llama" and requires a separate Meta licence above 700 million monthly active users. Calling a hosted endpoint moves those obligations onto the provider.
When is neither a serving engine nor an API the right answer?
Below a handful of concurrent users. For laptop development, air-gapped evaluation or a single-user demo, Ollama and llama.cpp are both MIT, start in one command and skip the configuration burden of a batching server you will get no throughput payoff from.
Sources
- vLLM repository, licence and supported hardware github.com
- vLLM releases github.com
- vLLM Ascend hardware plugin github.com
- LMDeploy repository, licence and build targets github.com
- LMDeploy Huawei Ascend getting started, supported Atlas devices lmdeploy.readthedocs.io
- LMDeploy releases github.com
- Groq model catalogue and per-token pricing console.groq.com
- Your Data in GroqCloud, storage location and retention console.groq.com
- Groq services agreement, export control terms console.groq.com
- Groq Helsinki data centre announcement groq.com
- DeepSeek API pricing, peak and off-peak windows api-docs.deepseek.com
- DeepSeek API rate limits and concurrency caps api-docs.deepseek.com
- DeepSeek privacy policy, data storage jurisdiction cdn.deepseek.com
- OpenAI supported countries and territories developers.openai.com
- Anthropic supported countries and regions anthropic.com
- Google Cloud machine learning service locations, africa-south1 docs.cloud.google.com
- Llama 4 Community License Agreement llama.com
- Qwen3-235B-A22B model card, Apache 2.0 huggingface.co
- The break-even on self-hosted open weights is 72 percent utilisation
Individual reviews: vLLM, LMDeploy, Groq, DeepSeek. All comparisons, or the full tool catalogue.