Downloading the weights saves only on research. Running the model is a cost of goods that climbs with revenue, and by that measure China's open challengers may not be cheaper to serve at all.

Kimi K3 spent a weekend rattling the AI industry. Moonshot's model — 2.8 trillion parameters, the largest set of open weights yet announced — landed close to the frontier on capability. Within days Alibaba answered with a preview of Qwen3.8 Max, and the argument on X ran hot. If a free, open-weight Chinese model can match the best American systems, what happens to the companies charging for the privilege?

Ben Thompson, who writes the technology newsletter Stratechery, took up that question in a column titled Who's Afraid of Chinese Models? His answer is that most of the alarm misreads the economics. Marginal cost is back, and it changes the whole shape of the argument.

What "free" actually refers to

Start with the word doing the damage. Download Kimi's weights (come July 27), skip the years and the hundreds of millions it took to train the thing, and the model has cost you nothing. True enough. But what you saved was research, and research is a fixed cost. Spend a million training a model, and you have spent that million whether you go on to earn a hundred thousand or a hundred million. The bill does not move with the business.

What moves with the business is COGS (the cost of goods sold), and for AI that cost is suddenly real in a way software forgot about years ago. Every answer a model returns burns compute. For most providers the inference bill tracks revenue almost one for one. Thompson's illustration is blunt: "if it costs 50 cents to generate the tokens that drive $1 in revenue, then $100 million in revenue will have $50 million in COGS." Open weights change none of that arithmetic. A model you host yourself still has to be served, and serving is where the money goes.

Kimi is instructive precisely because it is not cheap — Moonshot priced the API under GPT-5.6 Sol but squarely at Sonnet's standard rate, nowhere near the bargain bin.

Model

Input, per million tokens

Output, per million tokens

Kimi K3

$3

$15

GPT-5.6 Sol

$5

$30

Claude Sonnet 5

$3

$15

Sonnet 5 runs an introductory $2/$10 through August 31, 2026; $3/$15 is the standard rate that takes effect after.

The startup that could have undercut everyone chose frontier rates instead. That decision says more about the open-weight market than any benchmark.

A token is not a token

Even the sticker price flatters Kimi. Jensen Huang, Nvidia's chief executive, likes to call the machines his company builds "token factories," and for a chip that treats every model alike, the framing holds: tokens per second, tokens per watt, cost per token. Then reasoning broke the meter. Models now think before they answer, and the thinking is billed as output. Kimi reportedly spends far more tokens than Sol to reach the same result, which quietly swallows its per-token discount. Agents widen the gap again, since some models need many more steps to finish the same job.

So the fungible thing is the answer, never the token beneath it. If two models produce the same working code, the code is interchangeable; the tokens each burned to get there are not, and that difference drops straight into COGS. Intelligence becomes the commodity. And a commodity market does not reward high prices, because a rival can sell the identical output. It rewards a better cost structure.

Scarcity sets the price

That reframing is the test the Chinese threat has to pass. Today demand for frontier intelligence outruns supply, and the bottleneck is compute. The shortage lets Nvidia charge fat margins, lets its customers resell capacity to Anthropic at a markup, and lets Anthropic mark the tokens up once more on top. Buyers are looking at prices inflated by scarcity, not at the true marginal cost of serving a request. Under all that markup the Chinese weights read as a bargain, yet Thompson doubts they are genuinely cheaper to run.

Which invites the obvious question. If the labs are fine, why do they act so frightened? They are still anchored to a world where training dominated the budget, and steep inference prices were how you funded the next training run. That world is closing. Inference is set to outpace training, the agent wave is enormous, and a business that can make it up on volume can let prices slide once the compute lands. That is the old commodity playbook: lower prices, heavier usage, a cost structure below the marginal supplier's. It is not a Chinese ambush.

The illusion that was always an illusion

Underneath all this runs a longer point. Software taught a generation that serving one more customer costs essentially nothing. It never quite did. Even returning a static page carries a price, tiny as it is; assembling a social feed on the fly, or ranking a search result and preprocessing the request behind it, carries a good deal more.

Inference just makes the cost impossible to look past, and a few thousand extra requests now surface plainly on the invoice. Zero marginal cost was a serviceable approximation for two decades — AI is the moment it stops being close enough to true.

None of this makes Kimi K3 and its Chinese peers unimportant. It makes them ordinary, one more set of suppliers in a market that has finally started to behave like one.