Moonshot's model still ranks below Fable 5 and GPT-5.6 Sol overall, and its weights are promised to open by July 27.

Moonshot AI's Kimi K3 reached first place on the Frontend Code Arena leaderboard on July 16, 2026 with 1,679 points, passing Claude Fable 5 and lifting the model line from 18th, where its predecessor K2.6 sat. It came first in six of the seven frontend domains scored and second in gaming. It is the first Chinese model and the first open-source model to lead that human-voted ranking.

The model runs with context windows reaching one million tokens, and it takes text and images as input natively. Moonshot introduced it under the banner of open frontier intelligence, pointing to techniques it calls Kimi Delta Attention and Attention Residuals, and said the model is already live across Kimi.com, Kimi Work, Kimi Code, mobile apps and the company's API. The weights, it promised, will be opened by July 27, 2026.

A million input tokens on K3 costs $3 and a million output tokens $15, with cached input billed lower. OpenAI asks $5 and $30 for GPT-5.6 Sol; Anthropic asks roughly $10 and $50 for Claude Fable 5.

Across benchmarks taken together, K3 sits below both Fable 5 and GPT-5.6 Sol, trailing Fable 5 in several test categories, though it beats Claude Opus 4.8 and GPT-5.5 on coding and agent evaluations. Moonshot said the model still lags the strongest closed systems, which it named as Claude Fable 5 and GPT 5.6 Sol, while claiming frontier-level results throughout its own evaluation suite. The company also said the user experience remains noticeably short of what those two deliver.

Vercel chief executive Guillermo Rauch posted on X that this was the first time an open model had beaten every proprietary one on this broad web engineering benchmark, adding that benchmarks do not always tell the whole story. Wharton's Ethan Mollick called K3 the closest to the frontier yet, while likewise urging care in reading the figures. ProgramBench maintainer Ofir Press said Kimi leaned on a measure the benchmark's creators advise against — an average of how much of each program was implemented, rather than a count of programs that fully work — which can flatter a model. Artificial Analysis found K3's accuracy on AA-Omniscience rose to 46% from K2.6's 33%, while its hallucination rate worsened to 51% from 39%.

The release seems to have unsettled Wall Street: investors spent Friday selling chipmakers such as Nvidia and Intel, and the Nasdaq ended about 1% down.

Whether the scores hold up in other people's hands waits on July 27, when the weights and the full model code are due and developers can download K3, host it themselves and change it.