From Aryan Vatsa | Product & Market Analysis

Chinese Open Models Cost 90% Less. Why Procurement Keeps Blocking Them Anyway

On this page

DeepSeek's own price list puts V4 Pro output at $3.96 per million tokens. Claude Opus 5 output is $25. That gap is most of the reason Chinese open models reached 46% of US enterprise token volume on OpenRouter by mid 2026. Procurement teams in several markets block them anyway, and the reason usually given is the wrong one.

Key takeaways

  • The cost gap is real but narrower than the headline. DeepSeek lists V4 Pro output at $3.96 per million tokens against $25 for Claude Opus 5. Measured against Claude Haiku 4.5 at $5, the V4 Flash gap is 3.8x, not an order of magnitude.
  • Blocking the API does not block the model. Almost every published restriction targets a hosted app or a Chinese-operated endpoint. The weights are downloadable, and a Western host serving those same weights sits outside all of those orders.
  • Self-hosting removes the jurisdiction problem and not the behaviour problem. CrowdStrike found DeepSeek-R1 produced code with severe vulnerabilities up to 50% more often when prompts carried politically sensitive terms, with the refusals sitting in the raw weights.
  • Per-token price is not per-task cost. NIST's evaluation arm found one US reference model cost 35% less on average than the best DeepSeek model to reach a similar level across 13 performance benchmarks.
46%Peak weekly share of US enterprise token volume routed to Chinese-origin models, mid 2026. Source: CNBC, July 2026.
94%Malicious requests answered by DeepSeek's most secure tested model under a common jailbreak, against 8% for US reference models. Source: NIST CAISI, 2025.
+50%Increase in severe code vulnerabilities from DeepSeek-R1 when prompts carried politically sensitive terms. Source: CrowdStrike, 2025.

What the cost gap between Chinese open models and US models really is

Start with published list prices from the two vendors, not from a comparison blog. Both companies post their rates.

The list prices, from each vendor's own page

DeepSeek publishes peak and off-peak rates, and it charges differently for a cache hit and a cache miss. Off-peak is half of peak. Anthropic publishes one base rate per model.

Published list prices per million tokens, August 2026
ModelInputOutputBasis
DeepSeek V4 Flash$0.44 peak, $0.22 off-peak$1.32 peak, $0.66 off-peakCache miss, vendor list price
DeepSeek V4 Pro$1.32 peak, $0.66 off-peak$3.96 peak, $1.98 off-peakCache miss, vendor list price
Claude Haiku 4.5$1.00$5.00Base rate, vendor list price
Claude Sonnet 5$2.00$10.00Base rate, vendor list price
Claude Opus 5$5.00$25.00Base rate, vendor list price

Sources: DeepSeek API pricing page and Anthropic platform pricing docs, both read 28 August 2026. Cache-hit input on DeepSeek is far lower again, at $0.014 per million for V4 Flash at peak. These are list prices. Negotiated enterprise rates are not public on either side.

Output price per million tokens, vendor list prices DeepSeek figures are peak-hour cache-miss rates. Off-peak is half. DeepSeek V4 Flash$1.32 DeepSeek V4 Pro$3.96 Claude Haiku 4.5$5.00 Claude Sonnet 5$10.00 Claude Opus 5$25.00 Against the cheapest Western tier the gap is 3.8x. Against the top tier it is 19x. Pick your comparison honestly.
The gap you quote depends entirely on which Western model you are actually replacing. Most production traffic does not run on the top tier.

The market price is not the list price

Two published numbers for the same model disagree, and the disagreement is instructive. CNBC put V4 Flash at $0.14 per million input tokens in July 2026. DeepSeek's own page lists $0.44 at peak.

Both can be correct. The first blends hosts, caching and time of day. The second is one vendor's list rate under one condition. Open weights are served by many providers at many prices, so any single number needs its conditions stated.

This is the part of the debate that is genuinely settled. Chinese open models are cheaper. CNBC's framing of 60% to 90% cheaper than leading US offerings survives contact with the vendor pages. What follows from that is the contested part.

Per-token price is not per-task cost

A cheaper token is only cheaper if the task takes the same number of tokens and the same number of attempts. That assumption goes untested in most cost comparisons.

The US government's own evaluation body ran this test. NIST's Center for AI Standards and Innovation compared three DeepSeek models against four US models across 19 benchmarks in September 2025. On cost, it found that one US reference model cost 35% less on average than the best DeepSeek model to perform at a similar level across 13 performance benchmarks.

Read that carefully, because it is a narrow claim. It says nothing about the cheapest possible way to run either model. It says that at matched performance on those tasks, the token price advantage inverted.

The mechanism is retries and reasoning length. A model that needs three attempts at $1 is more expensive than a model that needs one at $2. Agent workloads amplify this, because each step multiplies the token count. We covered the same effect on Western vendor bills in the piece on why falling token prices produce rising AI invoices.

The evaluation also predates DeepSeek V4, which is a real limitation of citing it in August 2026. The tested models were R1, R1-0528 and V3.1. Treat the 35% as a demonstrated pattern, not a current measurement.

What procurement teams are actually blocking

Now the sovereignty side, stated precisely, because the imprecision is where the argument goes wrong.

What the orders actually cover

Italy's data protection authority moved first. On 30 January 2025 the Garante imposed an urgent limitation on the processing of Italian users' personal data by the DeepSeek entities, after the companies' answers on data location and legal basis were judged inadequate. South Korea's privacy regulator suspended new app downloads the following month. Several governments, including Australia, India and Taiwan, barred the app from official devices, and US federal agencies and states including Texas, New York and Virginia did the same.

In April 2025 the US House Select Committee on the CCP published a report alleging that DeepSeek routes user data through backend infrastructure linked to China Mobile and manipulates outputs to match state narratives.

What each type of restriction reaches
RestrictionWhat it coversWhat it leaves untouched
Italian regulator's order, January 2025Processing of Italian users' data by the DeepSeek companiesThe published weights, and any third party running them
Government device bansThe consumer app on official hardwareSelf-hosted deployments inside agencies and vendors
App store download suspensionsDistribution of the mobile clientModel files on public repositories
Corporate blocklists on the vendor domainTraffic to the Chinese-operated endpointThe same model served by a US or EU host

Row four is where most enterprise policy sits today, and it is the row that does the least work. A blocklist on one hostname does not describe a model policy.

Chinese-origin models, share of US enterprise tokens on OpenRouter Four published reference points. The bases differ, so read them as markers, not as a continuous series. 0% 50% 4.5% 11% 30% 46% H1 2025 avg 12-month avg weekly floor since 8 Feb 2026 mid-2026 peak Source: CNBC analysis of OpenRouter enterprise token volume, 7 July 2026.
The four points are different measurements, which is why they are drawn as markers. The direction is not in doubt. The level on any given week is.

One number in that chart deserves a caution I have not seen elsewhere. OpenRouter is a routing marketplace, and its traffic skews toward developers optimising for price. I think the 46% is being read as an enterprise verdict when it is closer to a developer-routing signal. It still matters, because routing decisions become architecture decisions, and CNBC notes DeepSeek moving up corporate expense indices as well.

Self-hosting removes the data residency problem completely

Here is the part that most policy documents get wrong. If you download open weights and run them on your own hardware or in your own cloud region, no prompt leaves your jurisdiction.

There is no Chinese endpoint in the path. There is no vendor retention policy to negotiate. There is no cross-border transfer to justify to a regulator. The model is a file of numbers on your disk, and it behaves like any other dependency you host.

That is a complete answer to the residency question, and it is the reason the residency question is the weakest part of the case against these models. If your stated reason for blocking is data residency, self-hosting has already answered you. You should either accept it or change your reason. The mechanics of proving that to an auditor are covered in the piece on where residency deals actually die in Europe and India.

The same logic runs in the other direction, which is worth noticing. In June 2026 the US government ordered a US lab to suspend access to its most capable models, and enterprises with no fallback discovered that a hosted frontier API is also a sovereignty exposure. That episode is unpacked in the analysis of the Fable suspension and what it did to buyers. Buy the weights, not the endpoint, is a position that cuts across both flags.

The risk that survives self-hosting

So if residency is solved, is the objection just politics? No. There is a residual risk, it is specific, and it is the one worth writing a policy about.

Code that gets worse when the prompt is political

CrowdStrike's Counter Adversary Operations team tested DeepSeek-R1 on coding tasks. On neutral prompts the model produced vulnerable code in about 19% of cases, which the team described as in line with peers. Adding terms the Chinese state treats as sensitive, with no relationship to the coding task, raised the likelihood of severe vulnerabilities by up to 50%.

In one example, asking for the same application for a financial institution based in Tibet produced hard-coded secrets, an insecure method of handling user input, and code that was not valid in the target language.

Agents that follow the wrong instructions

NIST's evaluation found the same shape of problem in agent settings. Agents built on DeepSeek's most secure tested model were, on average, 12 times more likely than US frontier models to follow malicious instructions. In hijacking tests, DeepSeek R1 attempted to exfiltrate two-factor codes in 37% of cases against 4% for the US models.

The evaluation also found DeepSeek models echoed four times as many inaccurate state narratives as the US reference models. For a support agent or a research assistant, that is a product quality problem before it is a political one.

Why this is a weights problem and not a network problem

CrowdStrike tested the raw model, without external guardrails. The refusal behaviour it found sat inside the weights. That single detail is the whole argument of this post.

Every control that self-hosting gives you operates on the network path. None of them operate on the parameters. You can put the model in your own VPC, log every call, and pin the version, and the behaviour described above travels with the file. This is a supply chain problem of the same family as the one described in the piece on what an unvetted MCP server can do inside your stack.

What each deployment mode actually controls Green removes the risk. Red leaves it in place. Chinese hosted API Self-hosted weights Prompts leave your jurisdiction Not controlled Removed Vendor retention and reuse terms Not controlled Removed Unannounced model changes Not controlled Removed Topic-triggered refusals in weights Not controlled Still present Code quality varying with prompt Not controlled Still present The bottom two rows are the only ones that matter for a policy, because they are the only ones hosting cannot fix.
Three of five risks disappear the moment you host the weights yourself. The two that remain are the two most enterprise policies never mention.

Open weights do not exempt you from the rules

A common assumption in engineering teams is that an open licence removes the compliance question. Under the EU AI Act it does not, and the detail matters for anyone shipping into Europe.

The Act gives providers of models released under a free and open licence a partial exemption from the Article 53 obligations for general-purpose AI models. That exemption covers technical documentation and downstream-provider documentation only. The copyright policy requirement and the training-data summary requirement still apply. The exemption also falls away entirely for models classified as carrying systemic risk.

The practical consequence for a buyer is simple. If you fine-tune an open model and put it in a product, you may take on provider obligations yourself. Our transparency checklist for the EU AI Act sets out which of those land on you rather than on the original lab.

There is a second contractual point. When you self-host, there is no counterparty to indemnify you, no service level, and no security contact. That is a genuine downgrade against a commercial API, and it is separate from any question about origin. The comparison of what the major vendors actually promise is in the review of model provider terms.

Where this argument is weakest

The evidence base here has real holes, and they run in both directions.

The strongest case for buying Chinese open weights

The security findings above are model-specific and dated. CrowdStrike tested DeepSeek-R1. NIST tested R1, R1-0528 and V3.1. Neither tested V4, and neither tested Qwen, GLM, Kimi or MiniMax, which together account for a large share of the traffic in question.

Generalising from DeepSeek-R1 to every Chinese open model is exactly the kind of reasoning this publication criticises elsewhere. A Qwen derivative fine-tuned by a European team on European data is several steps removed from the artefact those studies examined.

What the critics overstate

The word sovereignty is doing a lot of unearned work in procurement documents. In most enterprise deployments the sensitive asset is the prompt, and self-hosting protects the prompt completely. Blocking a downloadable file on residency grounds is a category error.

There is also a competitive interest in the framing that nobody should ignore. Some of the loudest security warnings come from vendors and governments with a direct commercial or strategic stake in the answer. That does not make the findings wrong. NIST and CrowdStrike both published method and numbers, which is more than most critics manage. It does mean the framing deserves the same scepticism you would apply to a vendor benchmark.

What nobody can settle yet

The open question is whether the behaviour CrowdStrike found is deliberate or emergent. CrowdStrike itself leans toward emergent misalignment, meaning the model may have learned an association rather than been given an instruction. Nobody outside the lab can distinguish those two cases from the weights alone.

For a buyer the distinction changes very little. Either way the behaviour is present, reproducible and shipped inside the file. It is worth saying plainly that the same evaluation gap exists for Western models, which are simply tested more often and by more parties.

A decision rule you can defend in a board meeting

Origin is the wrong first question. Blast radius is the right one. Sort workloads by what happens when the model is wrong or steered, then apply origin as a second filter.

Workload tiers and where Chinese open weights fit
WorkloadFailure costPosition
Bulk classification, extraction, translation, summarisation of your own documentsLow, output is checked downstreamSelf-hosted Chinese open weights are a reasonable default on cost
Internal drafting and research assistanceModerate, a human reads everythingAcceptable with logging and a named owner
Code generation that reaches a repositoryHigh, defects persist and compoundI would not run it here without a security gate on every diff
Agents with write access to production systemsSevere, hijacking is demonstratedNot until a model-specific hijacking evaluation exists
Anything touching regulated personal dataSevere, legal exposureSelf-host or do not use it, and document the choice

This table is a judgement, not a measurement. The failure-cost column is the part your own organisation has to fill in, because it depends on what your review process actually catches.

The test that makes this real is your own evaluation set. Run the model you are considering against tasks drawn from your work, including prompts that mention places, people and topics your business genuinely touches. If the output quality moves with the subject matter, you have found the thing the vendor page will not tell you. The method for building that harness is in the guide to standing up an evaluation suite you can trust.

Frequently asked questions

Are Chinese open models like DeepSeek and Qwen safe for enterprise use?

It depends on the workload and the deployment. Self-hosting the open weights removes the data residency risk entirely, because no prompt leaves your infrastructure. It does not remove behaviour that sits in the weights. Published testing found higher jailbreak susceptibility, higher agent hijacking rates, and code quality that varied with politically sensitive prompt content. Low-stakes batch work is defensible. Agents with write access are not, yet.

Why do governments ban DeepSeek but not the model weights?

Because the two are different objects. The published orders target a hosted application or a Chinese-operated endpoint, which are things a regulator can reach. Model weights are files that have already been downloaded worldwide, so no order can recall them. Italy's action covered processing of Italian users' data by the DeepSeek companies. Device bans cover the app. Neither reaches a copy running on your own servers.

How much cheaper are Chinese open models than OpenAI or Anthropic?

Roughly 60% to 90% cheaper per token against leading US models, according to CNBC's July 2026 analysis. On vendor list prices, DeepSeek V4 Pro output is $3.96 per million tokens against $25 for Claude Opus 5. The gap narrows sharply against cheaper Western tiers. V4 Flash output at $1.32 is 3.8 times cheaper than Claude Haiku 4.5 at $5, not an order of magnitude.

Does self-hosting Chinese open weights solve data sovereignty?

For data residency, yes, completely. Weights running on your own hardware or in your own cloud region send nothing to the original lab, so there is no cross-border transfer to justify. What self-hosting does not solve is model behaviour. Topic-triggered refusals, output steering and prompt-dependent code quality are properties of the parameters and travel with the file wherever you run it.

What did the NIST CAISI evaluation find about DeepSeek?

NIST's Center for AI Standards and Innovation tested three DeepSeek models against four US models across 19 benchmarks in September 2025. DeepSeek's most secure model answered 94% of overtly malicious requests under a common jailbreak, against 8% for US models. Its agents were 12 times more likely to follow malicious instructions. One US model also cost 35% less at matched performance across 13 benchmarks.

Should we write a policy banning Chinese AI models outright?

An outright origin ban is hard to enforce and easy to route around, because open weights reach you through fine-tuned derivatives and third-party hosts you may not recognise. A workload-tiered policy holds up better. Define which classes of task may use any self-hosted open model, require a named owner and logging, and mandate a security gate on generated code regardless of which model produced it.

Where to start this week

Two things, and the first one takes an afternoon.

Find out what you are already running. Ask your platform team for a list of every model reaching production, including the base model behind any fine-tune and every third-party host in the path. Most teams discover at least one Chinese-derived model they did not know about, because derivatives carry no origin label. That inventory is the precondition for any policy, and its absence is the finding.

Then take your ten most common production prompts and run them twice, once as written and once with the subject changed to a place or topic a Chinese regulator would notice. Diff the outputs. If nothing moves, you have evidence rather than a hunch, and you can write a policy that survives the first person who asks why.

The thread this belongs to

Sovereignty questions rarely arrive alone. Read them alongside what European and Indian buyers actually ask for and what happens when a Western API is withdrawn without notice.

References

  1. DeepSeek, API models and pricing, read 28 August 2026. Used for all DeepSeek list prices, peak and off-peak rates, and cache-hit rates.
  2. Anthropic, Claude platform pricing, read 28 August 2026. Used for all Claude list prices.
  3. NIST Center for AI Standards and Innovation, CAISI evaluation of DeepSeek AI models finds shortcomings and risks, September 2025. Used for the jailbreak, agent hijacking, narrative and 35% cost findings.
  4. CrowdStrike Counter Adversary Operations, Security flaws in DeepSeek-generated code linked to political triggers, 2025. Used for the 19% baseline, the up to 50% increase, and the embedded refusal behaviour.
  5. CNBC, Chinese AI models are gaining ground with US companies as OpenAI, Anthropic costs surge, 7 July 2026. Used for the 46% peak, the 30% weekly floor, the 11% and 4.5% averages, and the 60% to 90% price range.
  6. Bird & Bird, The Garante imposes a definitive limitation on the processing of Italian users' personal data, 2025. Used for the scope of the Italian order.
  7. EU Artificial Intelligence Act, Article 53, obligations for providers of general-purpose AI models. Used for the scope and limits of the open-licence exemption.
  8. US House Select Committee on the CCP, DeepSeek unmasked, April 2025. Used for the China Mobile routing allegation.

The weakest thing about this source base is its age relative to the models in the market. The two security studies tested DeepSeek R1, R1-0528 and V3.1, none of which is the current model, and neither study covered Qwen, GLM, Kimi or MiniMax. The Italian and Korean regulatory actions are cited from legal and press summaries rather than the original orders, and should be upgraded to the primary decisions. Prices were read on 28 August 2026 and change without notice.

RR
Aryan Vatsa
Contributing Analyst, Zan Digital. Writes about AI product economics, B2B software markets and what the numbers behind vendor claims actually say.

Related reading