An independent harness beat Anthropic's own tool on the same model in May; within two months the major labs had retuned theirs to favour their own models.
The decisive layer in AI systems has moved from the model to the software wrapped around it — the harness, which governs what a model can see, remember and do — and swapping that layer can shift performance more than swapping the model underneath.
That is the central claim of Mozilla's first State of Open Source AI report, drawn from a worldwide survey of more than 950 developers. Models themselves are becoming interchangeable, the report argues, so the real risk of power concentrating sits one tier up, in the harnesses.
The evidence it leans on comes from Terminal-Bench, a coding benchmark run in May, where a harness built by an outside party scored markedly higher than Anthropic's own tool — with both running on the same Anthropic model. Within two months, Anthropic and rival labs had rebuilt their tools so that each firm's harness now performs best with that firm's own model and worse with competitors'. The report reads this as a defensive barrier taking shape.
The report also points to cost pressure at two large engineering organisations. It says Microsoft cancelled most of its Claude Code licences in late June, after usage-based billing consumed the division's entire annual AI budget in a few months. Stripe, the report says, cut its AI operating costs by 73% by moving 50 million daily requests off proprietary vendor interfaces and onto open models running on infrastructure it owns and operates.
Trying open models, however, is not the same as running them. Some 79% of surveyed developers use them, but only 51% have taken them into production, against 63% for closed models. Álvaro Ruiz Cubero of SlashData, the firm that carried out the survey for Mozilla, says the gap is not purely a matter of model quality but of missing infrastructure, noting that deployment rates barely climb as companies grow — a sign, he says, of tooling and support that has yet to mature.
On raw capability, the report puts the best closed models 3.3 points ahead of the best open ones on the Chatbot Arena preference leaderboard, measured in March and down from more than eight points at the start of 2024. As TIME noted, the report itself concedes that the figure hides a jagged frontier, with closed systems still well ahead in significant areas. Anthropic's June release of Claude Fable 5 put closed labs back in front on the hardest reasoning and coding tasks.
The organisation's chief technology officer, Raffi Krikorian, told the magazine the report is partly advocacy for its campaign against concentrated power in the tech industry — and said no one should be surprised if it releases a harness of its own in a few months.