Can Cheap Chinese AI Models Rival OpenAI and Anthropic?

Are the New Cheaper Chinese AI Models Going to Rival the Current Big Players?

2026-08-02

Key Takeaways

  • They already do on volume. Chinese-origin models have taken at least 30% of the tokens US companies route through OpenRouter every week since 8 February 2026, peaking near 46%.
  • That figure averaged 11% over the preceding twelve months and stood at 4.5% in the first half of 2025.
  • Open-source Chinese models run roughly 60% to 90% cheaper than leading Anthropic and OpenAI offerings, according to OpenRouter’s own analytics team.
  • DeepSeek V4 Flash lists at $0.14 per million input tokens against $5.00 for GPT-5.5 — a gap of more than 30x at the extremes.
  • On capability, the best Chinese open-weight models still trail the strongest closed Western models on composite benchmarks, though the margin has narrowed faster than most forecasts predicted.
  • Revenue tells a different story to usage: Anthropic reportedly captures roughly half of total OpenRouter platform spend from a much smaller token share.
  • Moonshot’s Kimi K3 broke the pricing pattern at $3 input and $15 output per million tokens, signalling that Chinese labs will not compete on price forever.
  • The real constraints now are governance and supply risk rather than quality — both Washington and Beijing are moving to control cross-border model access.
Artificial Superintelligence (ASI) - artistic concept. Image credit: Alius Noreika / AI

Artificial Superintelligence (ASI) – artistic concept. Image credit: Alius Noreika / AI

The question of whether cheaper Chinese models will rival the established leaders has been overtaken by events. On measured developer usage they already rival them, and on several coding and agentic benchmarks they match or beat mid-tier Western releases. What they have not done is capture the revenue, the enterprise trust, or the top of the capability curve — and the fight over the next eighteen months will be decided there rather than on price.

The evidence for the usage claim is not anecdotal. A CNBC investigation published on 7 July 2026 found that the share of tokens US companies route through Chinese AI models on OpenRouter has stayed above 30% every week since 8 February, rising as high as 46%. Set against a 12-month average of 11% and a figure of 4.5% in the first half of 2025, that is one of the fastest supplier substitutions the software industry has seen.

The Price Gap Is the Whole Story — Almost

Justin Summerville, who works on data and analytics at OpenRouter, told CNBC that open-source Chinese models can run 60% to 90% cheaper than leading Anthropic and OpenAI models. The per-workload comparisons in that reporting were stark: roughly $4,811 through Anthropic’s Claude, $3,357 through ChatGPT, and $544 through Zhipu’s GLM for comparable work.

Model Origin Input price per million tokens Notes
DeepSeek V4 Flash China $0.14 284B total parameters, 13B active
GLM-5.2 China (Z.ai) Low-cost tier Led open-weight intelligence rankings in mid-2026
Kimi K3 China (Moonshot) $3.00 ($0.30 cached) 2.8T parameters, 1M context, $15 output
GPT-5.5 United States $5.00 Frontier closed model
GPT-5.2 United States $1.75 Mid-tier closed model

Harpreet Arora, head of agentic infrastructure at Vercel, put the mechanism plainly to CNBC: price is doing the work, and when a task does not need the best model, teams route it to the cheapest one that is good enough. Vercel saw GLM 5.2’s daily token volume grow roughly 27-fold in its first full week after launch, with customer count up around 80-fold — the fastest adoption of any model the platform tracked in 2026.

Named adopters have followed. Lindy moved its traffic from Claude to DeepSeek, reporting cost savings it described in the millions. DoorDash and Airbnb have both taken on Chinese models as cheaper alternatives, and Coinbase reportedly runs a large fleet of agents on them.

Why the Prices Can Be That Low

Three structural forces sit under the gap, and none of them are temporary.

The first is open weights. DeepSeek, Qwen, GLM and the Kimi line ship as downloadable models. Once anyone can host a model, no single vendor controls its price, and inference becomes a commodity trending toward the cost of the hardware. Closed Western flagships have no such pressure — their makers set the rate.

The second is efficiency engineering under constraint. Export controls pushed Chinese labs toward architectural savings rather than brute compute. DeepSeek’s sparse attention work, and Moonshot’s Mixture-of-Experts designs that fire under 2% of available experts per token, are direct products of that pressure. Some of this now runs on domestic silicon: GLM-5 and DeepSeek V4 are trained and served on Huawei Ascend hardware, which for Chinese state buyers is worth more than a benchmark point.

The third is deliberate strategy. Cheap tokens are a customer-acquisition tool for labs competing against incumbents with a head start, in much the same way that DeepSeek’s original low-cost releases bought global attention that advertising could not.

The Capability Picture Is More Even Than It Was

Independent leaderboards through 2026 have consistently shown the same shape: Chinese open-weight models occupying most of the top open positions, trailing the strongest closed Western models by a single-digit margin on composite scores, and beating them outright on specific tasks.

Kimi K2.5 came close to top proprietary systems on early benchmarks at roughly one-seventh the price of Claude Opus. Kimi K2.5 outperformed Claude Opus 4.5 on BrowseComp by a wide margin. GLM-5’s makers reported it surpassing Gemini 3.0 Pro on agentic coding. And Kimi K3 opened at first place on Arena’s Frontend Code leaderboard in blind developer testing, ahead of Claude Fable 5 — while Moonshot itself conceded K3 trails Fable 5 and GPT-5.6 Sol overall.

That last detail is the honest summary. The very top of the capability curve remains Western. Everything below it is contested, and readers weighing specific matchups will find the differences narrower than headlines suggest — a comparison that is easier to make now that mid-tier frontier models such as Claude Sonnet 5 are priced against Chinese flagships rather than above them.

Where the Rivalry Actually Breaks Down

Usage share is not revenue share. Anthropic reportedly takes roughly half of all OpenRouter platform spending from around 12% of token volume, because premium models are used for the tasks that justify premium prices. A model can win the majority of calls and still lose the majority of the market by value.

Nor is the arrangement free of risk. In June 2026 the US Department of Commerce ordered Anthropic to suspend access to Claude Fable 5 and Mythos 5 under export-control directives, restrictions lifted on 1 July — the first known instance of a government forcing a global model takedown, and a precedent that cuts in every direction. Reuters reported on 7 July that China’s Ministry of Commerce was in discussions with Alibaba, ByteDance and Z.ai about restricting overseas access to their most advanced models. US lawmakers opened inquiries into Airbnb and Anysphere over their use of Chinese models the following day. Trade policy is now a live input to model selection, much as it became for hardware procurement during the tariff cycle.

There is a data-governance dimension too. Because the leading Chinese models ship open weights, the mitigation is straightforward: self-host, or use a US or EU cloud provider, which preserves most of the cost advantage while removing the data-location question. Separately, research shared with Semafor in July 2026 found that models distilled from open Chinese checkpoints do not reliably reproduce the censorship behaviour of the originals — undercutting one of Washington’s stated concerns.

What to Watch

Three signals will settle this. First, whether Chinese labs hold their price advantage as they move upmarket — Kimi K3’s $3 input rate suggests some will not. Second, whether reasoning and long-horizon agent work, the last clear Western premium, stays that way; growth in high-end Chinese reasoning models was running near 30% weekly in mid-2026. Third, whether either government narrows the field by decree before the market resolves it.

For buyers today, the practical answer is a routing question rather than a loyalty one. Send the tasks that need frontier reasoning to a frontier model, and send everything else to the cheapest capable option. That is precisely what the OpenRouter numbers describe teams already doing.

If you are interested in this topic, we suggest you check our articles:

Sources: CNBC, MIT Technology Review, Tom’s Hardware, CNBC (DeepSeek V4), eWeek, Anthropic

Written by Alius Noreika

Are the New Cheaper Chinese AI Models Going to Rival the Current Big Players?
We use cookies and other technologies to ensure that we give you the best experience on our website. If you continue to use this site we will assume that you are happy with it..
Privacy policy