In the beginning, software was open. The source code for the first mainframe operating systems, as but one example, was shared with customers liberally because the concept of commercializing it rather than the hardware it ran on wouldn’t be invented for over a decade.
Models, likewise, were open before anyone thought about closing them for commercial purposes. As far back as 2012 AlexNet was distributed with open weights – a term, importantly, that is distinct from open source – and in 2014 Caffe created a proto-Hugging Face for open weight models with its Zoo. Up the evolutionary ladder, more recent examples like ELMo (Feb 2018) and GPT-1 (June 2018) were distributed with weights.
It wasn’t until nearly a year after the latter even that OpenAI’s successor to GPT-1, GPT-2, was initially withheld and kept closed. Multiple reasons were provided for this change: concerns that it would flood the internet with spam, concerns about the potential for offensive, racist output and so on. Nine months later, OpenAI relented and opened the closed model by providing weights.
As commercial vendors began to perceive the possibility of unprecedented product growth and commensurate financial returns, much as with software once upon a time, closed became an increasingly common choice from model makers. Not, or at least not just, for theoretical safety concerns, but as means of creating a scarcity commercial vendors could mine. In the United States, at least, the best performing models today are all closed. But as has been discussed previously, the gap between closed models and their open weight alternatives has been closing.
Their lengthy history notwithstanding, the real race between open weight and their closed model counterparts arguably began on Christmas in 2024. That, at least, was the moment when the market began to perceive the potential financial impact of open weight models. On that December day, a little known Chinese model company called DeepSeek released weights for an impressive model. Thirty-three days after that, the market – if only temporarily – shaved over half a trillion dollars off of NVIDIA’s market cap.
The stakes, in other words, were high, and the increasing parity in functional capabilities between closed and open models is making them that much more so.
But how should artifacts as complex as models be considered? What are the dimensions that will shape this race moving forward? There are many angles to consider, but the following are arguably the most important.
Geography
With rare exceptions like Inkling (Thinking Machines) and Nemotron (NVIDIA), the largest and best performing open weight models in 2026 were all Chinese in origin. Consumption followed the rapid improvement in China’s open weight models, with one report claiming that they represented 4.5% of enterprise token usage on OpenRouter in early 2025 only to jump that to 63% a month ago.
This geographical asymmetry in the popularity and support for open models has strategic implications that are both direct (traditional nation state competition) and indirect (long term developer affinity).
Inevitably, these tensions have led to political jockeying. Armed with escalating if unsubstantiated accusations of model distillation on China’s part, in late July, less than a week after Kimi K3 was released, the White House, via its Treasury Secretary and Science Advisor, floated a ban on US companies using Chinese models, sanctions and other punitive measures.
The response from the technology industry was immediate. Within two days, a trade association of small startups argued against a ban. Two days after that, NVIDIA’s Jensen Huang posted his first ever Tweet, a letter essentially arguing against restrictions on open models which was also signed by a wide range of companies and foundations like Dell, Hugging Face, IBM, the Linux Foundation, Microsoft, Mistral and Y Combinator. Notably absent were Anthropic, Google and OpenAI, but the latter two joined a day later and by the end of the week 235 companies were signatories.
The market, in other words, demands access to open weight models, and at least according to benchmarks the best open weight models in the world at present come from China. Geopolitical tensions, therefore, are inevitable.
Licensing
In mid-August, as had been rumored for some time, the Linux Foundation submitted its OpenMDW-1.1 for approval to the OSI. The license was originally created to make up for the perceived shortcomings of licenses built for software, not projects of which software was but one small part. It already governs 19 NVIDIA models as well as others from BAAI, IBM and Poolside. It is not clear, however, that it will be approved as an open source license – the thread discussing the license’s merits and shortcomings has reached over 90 messages, and concerns about triggers and definitions may yet sink its candidacy.
Beyond the difficulties of how or whether the term open source could and should be applied to models, weights pose further complications in that it is not a matter of settled law that copyright can be applied to them. And if it can’t, the open source licenses based on that legal concept cannot be used. This is why distillation, discussed in more detail below, is generally objected to on the violation of terms of service basis rather than copyright infringement.
Even if complicated questions of licensing are set aside, it would appear that the direction of travel for open and closed – for the industry and, curiously, for many model providers – is not clear. Broadly, of the 96 models RedMonk is currently tracking, 51% are open in some fashion with the remainder closed. Even within specific model authors, the release terms have changed from version to version. OpenAI is the canonical example here, from the (eventually) open GPT-2 (2019) to the closed GPT-3 (2020) to the open gpt-oss (2025). Meta likewise went from the open weight Llama to the closed Muse Spark, which was subsequently promised to be opened (but hasn’t as of publishing time). Moonshot went from an open weights K2 (but importantly not open source, MIT license branding notwithstanding) license to a more restrictive license for K3. Alibaba, for its part, did the exact opposite with Qwen 3.8-Max, opening the previously closed model at one size level (27B) and placing it under a standard Apache open source software license. Its larger 2.4T cousin, however, is under a commercially restrictive source available alternative.
Licensing, in sum, has a complicated future ahead of it with weights.
Capabilities
On the one hand, even with the incredible advances in open models this summer from the likes of Kimi K3, Qwen 3.8-Max and GLM 5.3, several benchmarks still give the closed frontier models a clear advantage. SWE-bench Verified, for example, place OpenAI’s Sol first in its ability to code, Anthropic’s Opus 4.8 second and K3 third with a ~12% gap between the latter two. In other more practical agentic benchmarks like the Vals Index v2, the delta between open and closed is even more stark: Opus 5 leads, while K3 is tenth. Throw in issues with hallucination rates and speed, and objectively the closed models still have a lead.
But the important question from a capability standpoint isn’t necessarily raw capability, but a threshold. The industry crossed an important inflection point in November of 2025 with first Opus and then Chat-GPT. Those models introduced an entirely new level of capability, making them massively more useful than the models that preceded them. The important question, then, has been not when or if there will be true parity between open and closed models, but when open weight models will match those original Opus capabilities: when will they be “good enough?”
The answer, it would appear, was a month ago. Or April, depending on your definition of good enough.
On the Artificial Analysis Intelligence Index, Kimi K3 was behind Fable 5 and GPT-5.6 Sol but comparable to Opus 4.8 and Chat-GPT 5.5. While the former models are more capable than the latter, Opus 4.8 and Chat-GPT 5.5 are incredibly capable, and the difference between them and their Fable and Sol cousins is for many tasks unimportant. There are similar results across other benchmarks, and in longer-running, agentic benchmarks like Elo, Program Bench and SWE Marathon, open weight models often place first.
There are some crucial infrastructure limitations at present – K3’s size in particular means it’s impractical to run locally. But the economics are increasingly compelling, as it’s roughly half the cost of Opus 4.8 and a third that of Fable 5.
The simple takeaway, however, is that open weight models already have the level of capability that changed the industry, and they’re only getting better.
Risks
While there is clearly a commercial incentive at work, Anthropic has consistently insisted that its primary concern with open weight models is not that they’re competitive, but the risks inherent to not having guardrails. The company famously declined to make its Mythos release publicly available out of concerns that it would make the exposure and exploitation of vulnerabilities by would-be attackers trivial. It also has by all accounts deep conviction that the models represent clear and present dangers in other, non-software technical arenas like biotech or nuclear.
While these latter concerns, as astute observers have noted, may be somewhat overstated given the substantial practical barriers inherent to designing biological or nuclear weapons versus attacking mere software, even if it is only merely that the cost of creating malware goes to zero, the models clearly bring risks with their respective benefits.
But just as models can be used for offense, so too can they be used defensively to protect from those attacks. While it was being assaulted by rogue OpenAI elements, notably, Hugging Face tried to turn to frontier models for assistance. The guardrails built-in to those, however, meant that the models declined to assist in the companies defense. Instead, Hugging Face had to turn to an open weight model without restrictions – GLM – to defend itself.
It’s not clear, however, whether or not the academic or theoretical concerns about risks are relevant. Even measures up to regulations such as the bans and sanctions mentioned above, models are software and software is inherently attainable – if not always easily stood up and operated in some cases.
Open weight models, in other words, are not a genie that can be put back in a bottle.
Economics
During a Q&A session with analysts during the 2025 reInvent conference, AWS CEO Matt Garman had a pointed observation about the economics of open weight models, noting that they are both costly to develop – even post-Deepseek – and in many cases lack an obvious economic model to reliably sustain their development. In short: “If I spent billions to build it, I wouldn’t give it away.”
The impact of economics on open weight model development has some fairly clear patterns. In a real sense, like open source before it, open weights are a tactic, not a business model.
- Market Creation: IBM and NVIDIA, among others, have clear economic incentives to build and release open weight models. The former can sell a variety of consulting and software services to offset their model development costs, and an IBM client relying on an IBM Granite model is more directly attached than one relying on a third party offering. NVIDIA, meanwhile, simply wants to sell as much hardware as possible, and to the extent that an open model like Nemotron helps do that, its margins can cover the development costs. Alibaba and Google, for their part, have been more aggressive with model development than cloud counterparts like Amazon and Microsoft, in part to use one area of strength to drive business to their respective cloud business units.
-
Subsidization: In addition to alignments with other internal core businesses, some model providers such as Google and Meta, have high margin businesses – ads, in both cases – that can help provide an economic footing for development of potentially complementary technology such as LLMs.
-
Venture Funding: Much of the closed US frontier market and to a lesser extent the Chinese open weight markets are venture bets on Dutch East India company-like projected valuations. To hit their nearly $1T value targets, Anthropic and OpenAI would have to sustain their meteoric growth, but at a price point with a profit margin rather than one subsidizing usage – all while simultaneously fending off competition from lower cost models. Providers of open weight models, both here in the US and from China, will face increasing questions regarding the economics of making weights available. If the economic picture is uncertain for frontier models which enjoy the artificial scarcity of closed software models, it’s even more so for open weight startups like Deepseek and Moonshot.
-
Nation State Support: That being said, it’s important to observe that China, unlike the US, has clearly identified open weight models as a means for achieving its decade-plus goal of AI superiority by 2030 and funded it as such. Through vehicles like the $8B National AI Industry Investment Fund, China is helping to subsidize companies like Alibaba, Deepseek, Moonshot and Z.ai. How far that support will extend as the costs pile up remains to be seen, but certainly the infusion of capital has greatly advanced Chinese model capabilities to date. One additional note from this section: China specifically has been accused of using open weight models as a form of price dumping, but definitionally this does not seem to apply any more than it does to open source.
The uncertain economics of open weight models may account for some regressions of models from open to closed licenses discussed above.
Release Timing
Another trend worth monitoring with respect to open weight models is the delay between when an open weight model is announced or the model made available, and when the weights are actually opened.
In the early days of open weight model development, delays in the release of open weights were largely if not entirely a function of safety reviews. These days it’s much more likely to be commercially motivated. And the delays overall seem to be getting longer.
Meta has gone from a same day release (Llama 2/3B) to 96 days (3.1) to indefinite as Behemoth’s weights were never released. Llama’s successor Muse, meanwhile, has had a complicated history with releases: Muse Spark was announced this spring, closed and described as “not suitable for open sourcing.” In early August, however, the company announced the availability of Muse Glimmer, a smaller, Apache-licensed distilled version of Spark with weights “coming soon.” Before those weights were made available, however, the company announced a newer version of the model whose weights would also be available “soon.”
In general, however, the release timing of open weight models seems to have become a function of the aforementioned geography. Chinese model weights are typically being provided on a 10-14 day timeframe; US models have taken frequently a hundred or more, if they are released at all.
Just as it will be important to watch if more open weight models become closed, however, it will be instructive to observe whether or not the timeframe for releases of Chinese open weight models stretches out.
Size
Just as open weight models have demonstrated a tendency to stratify by release date according to where they were built, the same is true of model size. While there are exceptions on both sides, in general Chinese open weight models (e.g. GLM, Kimi, Qwen) are typically larger – 70B and up – while their US counterparts (e.g. gpt-oss, Glimmer, Gemma, Granite) are smaller and more easily run on local hardware.
There are many reasons for this ranging from development costs to the models’ intended purpose. Kimi K3, for example, is designed to be directly competitive with the best frontier models, and its massive 2.8T parameter size – which means it is very unlikely to be run locally – reflects that need for broad scope capabilities. IBM’s much smaller 30B and under Granite models, however, are intended to be deployed and run in many cases on-prem for specific, tailored enterprise use cases.
It will be interesting to see over time whether this pattern holds, or whether we see more parity in model size across different geographies.
Distillation
Model distillation is essentially training one model on another model rather than novel datasets. One model asks a second model millions of questions, and uses the answers as input to improve itself. It’s a way of at once accelerating a model’s development while potentially obviating the need for massive amounts of new training data.
There are benign uses of distillation, and in fact many model providers distill their own models: Google’s Gemma is distilled from Gemini, Meta’s Glimmer is distilled from Spark, Anthropic’s Haiku from Opus and so on. But distilling a competitor’s model is at best frowned upon and at worst illegal. As mentioned briefly above, distillation can be an infringement of a model’s terms of service. Many models carry generally open licenses, but which restrict model training and distillation as a competitive defense.
Distillation is interesting, in that it reads on every other key theme above.
- Geography: material claims of distillation almost always originate in the US, aimed at competitive Chinese open weight models.
- Licensing: distillation is in a legal gray area with respect to copyright, forcing model providers to resort to clunky terms of service use restrictions in their licenses.
- Capabilities: if the claims of US providers are substantiated, improvements in open weight model capabilities could be, in part, a function of distillation.
- Risks: one common concern is that if open weight models can borrow from closed frontier alternatives, they could represent a similar threat but with none of the guardrails.
- Economics: given the incredible cost of training models, distillation at any real scale would have potentially massive economic implications for competing models.
- Release Timing: at least some of the growing reticence of US model providers to release their weights is attributable to fears of distillation.
- Size: as evidenced in the benign usage of distillation to shortcut the process of developing their own smaller models, distillation can be used to improve the ability and intelligence of smaller models.
The Net
As with every other AI artifact, open weight models are both infinitely more complex than can be captured through a few simple lenses and changeable by the day. Given that, the above is at best a basic rubric by which to judge open weight models and likely to be of transient value, particularly in the long term.
Dynamic as they are, however, as the industry seeks to better understand the role of open weight models, these macro considerations may be of use in appreciating their impact today and potential impacts moving forward.
Disclosure: Amazon, Google, IBM, Microsoft are RedMonk customers. Alibaba, Anthropic, Deepseek, Meta, Moonshot, the Linux Foundation, NVIDIA, OpenAI, the OSI, Poolside, the Thinking Foundation and Z.ai are not currently customers.





No Comments