Mozilla’s updated State of Open Source AI report makes a consequential, if qualified, argument for AI buyers: leading Chinese open-weight models have become close enough to the best closed U.S. systems that organizations should choose models by workload, rather than standardize on a single model category. The report’s fitted estimate, based on METR task-horizon data, places the gap between leading open and closed models at about 4.4 months. Mozilla published version 1.1 on September 15, with data current through September 1, according to reporting by Ars Technica and Tom’s Hardware.

From a model choice to a workload decision

That finding does not establish that open-weight models now match frontier systems in general. It does, however, challenge the assumption that a frontier API is automatically the rational default for every summarization, coding, extraction, classification, customer-support or internal-assistant task. When a near-frontier alternative costs materially less, the relevant question shifts from “Which model is best?” to “How much additional quality, reliability or task coverage does this particular workflow gain from the premium model?”

Mozilla’s comparison illustrates why that distinction matters. Ars Technica reported that Moonshot AI’s Kimi K3 was three points behind Anthropic’s Fable 5 on the Artificial Analysis Intelligence Index while listed at 30% of its cost. In another benchmark comparison highlighted by Ars, Z.ai’s GLM-5.2 scored within one point of Claude Opus 4.7 and 4.8 on Vals AI’s Terminal-Bench 2.1 while costing about one-fifth as much per completed task. Those figures are not interchangeable measures of universal capability, but they point to an economic reality that procurement teams cannot ignore: a small measured performance difference can coexist with a much larger difference in operating cost.

The most immediate enterprise consequence is likely to be model routing. Rather than selecting one vendor and applying it across every task, teams can reserve closed frontier models for work where their advantage is material and use open-weight options for high-volume, more bounded workloads. Mozilla CTO Raffi Krikorian told Ars that expert professional work, high-intensity retrieval and long-context tasks are categories where closed models still earn a premium. Ars also reported his view that either type can address tasks below roughly eight hours, while the eight-to-12-hour range is where closed frontier models retain the relevant edge.

This should be understood as a portfolio decision, not a declaration of independence from proprietary AI. A low-cost model used for routine work can lower total inference spending and make it easier to offer AI features more broadly. But workflows involving difficult reasoning, sensitive retrieval, long documents or a high cost of failure may still justify paying for the strongest available hosted system. DoorDash, according to Ars, provides an example of that division: it uses Kimi for routine work and reserves Fable for harder tasks that would take human experts longer to complete.

What the task-horizon estimate does—and does not—show

The report’s use of task-horizon data also deserves careful interpretation. METR defines a 50% task horizon as the human-expert task duration at which an agent is predicted to succeed half the time, based on a fitted relationship between task duration and success. Its published suite is primarily made up of self-contained software engineering, machine-learning and cybersecurity tasks—not the full range of enterprise work. METR explicitly cautions that such measures should not be read as a claim that an AI can autonomously perform all work of an equivalent duration, particularly when real jobs depend on organizational context, tacit knowledge and human interaction. (metr.org)

That limitation is central to Mozilla’s “months behind” framing. A 4.4-month fitted gap can be useful as a directional comparison, but it is not a service-level agreement. Model performance remains uneven across domains and can change substantially with the agent harness, tool access, context design, retrieval quality and evaluation method. METR says its measurements can be delayed or incomplete, and it warns that measurements above 16 hours are unreliable with its current task suite. (metr.org) For enterprise buyers, internal evaluations on representative tasks therefore remain more important than a leaderboard position or a single aggregate estimate.

Costs, deployment and openness

The apparent cost advantage requires similar discipline. Mozilla’s reported comparisons are hosted API-to-API list-price comparisons, not a guarantee that self-hosting will be cheaper in every environment. Tom’s Hardware noted that serving a leading open model can carry substantial infrastructure requirements, citing a Mozilla configuration for Kimi K3 with 64 or more accelerators. Open weights provide the option to download major model components, but they do not erase the costs of hardware, capacity planning, inference optimization, reliability engineering, security review, monitoring and support.

There is also an important distinction between open-weight availability and fully open-source development. The report says developers commonly withhold training data, data-pipeline details and training code, and Tom’s Hardware reported that none of the 16 notable releases covered met the Open Source Initiative’s definition requiring the data recipe. That means open weights can offer deployment control and vendor flexibility without necessarily delivering full reproducibility or transparent provenance. For regulated or security-sensitive deployments, the location of a model’s developer, the quality of its documentation, licensing terms, update cadence and security processes remain part of the buying decision.

Why commercial adoption may diverge from usage

Mozilla’s usage and revenue observations underline why capability convergence may not immediately translate into commercial convergence. The report found that eight of OpenRouter’s 10 most-used models by token volume in August were open-weight, with seven of those eight developed in China. Yet Linux Foundation data cited by the report showed closed models taking 96% of model-layer revenue on OpenRouter from May through September 2025. The contrast suggests that open models can win substantial volume while premium closed models continue to capture higher-value spending.

The report’s larger contribution, then, is to make “frontier” a narrower purchasing category. The premium is most defensible where better results alter business outcomes: difficult expert tasks, demanding retrieval, long-context work and tasks for which failure or manual remediation is expensive. Everywhere else, a capable open-weight model may increasingly be the baseline against which proprietary offerings must justify their cost. Organizations that build evaluation, routing and governance around that distinction will be better positioned than those treating the open-versus-closed choice as an all-or-nothing technology bet.