Chinese open-weight models are no longer a curiosity for US developers. The price gap is large enough that real workloads are moving, even when the politics are ugly.
If you build with AI at scale, you already know the quiet panic behind this story. The bill arrives before the strategy does. Coinbase has now put a public number on it: The Information reported that CEO Brian Armstrong said the company cut AI spending nearly in half while token use kept rising, helped by cheaper defaults including Z.ai's GLM-5.2 and Moonshot AI's Kimi K2.7 Code.
That's the real shift. Not every task needs the most expensive model in the stack. Code review, summarization, internal drafting and routine agent work can burn through millions of tokens without needing Anthropic's strongest system or OpenAI's best model every time. Once teams can route those jobs elsewhere, the old default starts to look lazy.
The pricing explains why developers are moving. OpenRouter lists Claude Opus 4.8 with a weighted average input price of $1.75 per million tokens and an output price of $25 per million tokens. Vercel's AI Gateway lists GLM-5.2 at $1.40 per million input tokens and $4.40 per million output tokens through several providers, with cheaper cached reads. That gap isn't theoretical. It shows up in the budget meeting.
The routing layer is deciding
Z.ai released GLM-5.2 on June 16 with a one-million-token context window and a coding focus. CNBC reported that Harpreet Arora, Vercel's head of agentic infrastructure, said daily token volume for GLM-5.2 grew about 27 times in its first full week on Vercel, while the number of customers using it rose about 80 times. His blunt read was right: "Price is doing the work here."
Vercel CEO Guillermo Rauch gave the model another kind of signal. In a public LinkedIn post, he said he was "genuinely impressed, almost shocked" by GLM-5.2's coding ability and added, "This changes things." You don't have to treat a CEO's post as a benchmark. You should treat it as evidence that the model was good enough to make serious developer-tool people stop and test it.
Then Kimi K3 arrived on July 16. Moonshot AI's model has 2.8 trillion parameters, according to AP and other reports, and its full weights were scheduled for release by July 27. Within days, demand for the hosted product pushed Moonshot close to its compute limits. Reuters reported on July 20 that the company temporarily paused new subscriptions after the launch strained capacity.
That's not hype. That's an operational bottleneck.
The aggregate traffic tells the same story. AP reported that, based on OpenRouter data over the past month, the five most popular models on the platform were Chinese. Separate CNBC-linked reporting found Chinese models holding more than 30 percent of OpenRouter token share every week since February 8, with peaks around 46 percent, compared with 4.5 percent in the first half of 2025. Developers can talk about national champions all day. Their API calls are voting on price.
Washington can slow it
The policy fight is now chasing the engineering decision. AP reported that the Trump administration accused Moonshot of using covert methods to build K3 from Anthropic's Fable, while Chinese officials rejected distillation allegations as groundless. Treasury Secretary Scott Bessent has also warned that sanctions could follow if Chinese firms are found to have stolen US technology.
Security teams aren't wrong to care. If your code, customer messages or internal documents flow through a hosted Chinese API, you have a data-governance issue to answer before anyone celebrates the savings. Self-hosting open weights shifts the risk profile - but it doesn't make inference free. Think about what it actually takes to run a 2.8 trillion-parameter model: serious hardware, sustained power draw, and the kind of operational discipline that most startups simply don't have. Not a back-room project.
US models still have the high ground on full-range capability, and that gap matters. AP quoted Arena co-founder and CEO Anastasios Angelopoulos saying Chinese models still trail American leaders across overall capability - and that gap matters for hard reasoning, sensitive enterprise work and tasks where a bad answer is expensive. But AI spending isn't monolithic. A huge share of token volume goes to work where good enough is exactly good enough.
Look, OpenAI and Anthropic aren't being abandoned. They're being removed from the default path. Developers are benchmarking, routing selectively and keeping the expensive model for the jobs that earn it.
That's how the market changes first.
Also read: Data centers have been cutting your electricity bill for years and the AI buildout may end that • Baseten built the fastest GLM-5.2 API on earth and the playbook tells you where inference is heading • AMD delivered better stock returns than Nvidia in H1 2026 but the AI chip war is far from over