Why Open Weight AI Is Gaining Ground as US Models Get Locked Down

Why Open Weight AI Is Gaining Ground as US Model Access Becomes Politicised

👁9views

Open weight models are winning because teams can run them on infrastructure they control, tune guardrails to real use cases, and avoid blanket refusals that block legitimate work like security incident response. Meanwhile locked down commercial APIs increasingly treat sensitive but valid queries as threats, pushing security teams, researchers, and cost sensitive builders toward transparent, self hosted alternatives like Kimi K3 and GLM 5.2.

CloudScale AI SEO: Article Summary
  • 1.
    What it is
    Open weight AI models like Kimi K3 and GLM 5.2 are overtaking closed US frontier models in enterprise use, driven by weak commercial moats and new export control actions. The article traces how Chinese open weight adoption jumped from under 2 percent to roughly 61 percent of top model traffic on OpenRouter by mid 2026.
  • 2.
    Why it matters
    Enterprises are turning to open weight models because closed API guardrails block legitimate work, such as incident response teams analysing real attack payloads, while giving them infrastructure control that survives sudden regulatory shutdowns.
  • 3.
    Key takeaway
    The same US government penalised Anthropic for refusing to loosen guardrails and separately used export control law to yank two other closed models offline for being too easily jailbroken, an incoherent policy that is pushing users toward open weight alternatives.
~13 min read

Two screenshots crossed my phone within about a day of each other this week, and together they say more about where this industry is heading than most of the analyst notes I read.

The first is Moonshot AI announcing it had to pause new subscriptions to Kimi K3 because demand pushed its GPUs close to their limits within 48 hours of launch. The second is from a colleague on a security team, describing why they could not use frontier commercial APIs for a piece of incident response work. The safety guardrails on those APIs could not tell an incident responder analysing real attack payloads from an actual attacker submitting them, so every serious request got blocked. The team ran the forensic analysis on GLM 5.2 instead, an open weight model. The weights are the billions of numerical parameters a model learns during training. They are what the model actually consists of: feed it a prompt, and the weights are what determine, step by step, which word or token comes next. Everything the model appears to “know,” every pattern of language, code, or reasoning it has picked up, lives in those numbers rather than in any separate database or rulebook. Open weight means those parameters are published for anyone to download, run, and modify, rather than sitting locked inside a vendor’s servers and reachable only through an API call. That let the team run GLM 5.2 on infrastructure they controlled.

Neither of those is a headline on its own. Put them next to Ben Werdmuller’s recent piece on werd.io, this month’s AI stock slump, and the export control saga I wrote about last month, and a pattern comes into view that every enterprise technology leader should be paying attention to.

1. The moat that was not there

Werdmuller’s argument is worth engaging with directly. His claim is that frontier models, as standalone products, have very little moat beyond brand loyalty and switching costs. The real moat sits in the enterprise layer around the model: contracts, integrations, and the operational features that make a model usable at scale. At the protocol level, switching providers can be as simple as changing an endpoint. In production it still means rerunning evaluations, adjusting prompts, redoing security review, and checking that tool calling and refusal behaviour have not shifted underneath you, so call the switching cost manageable rather than minimal. Even accounting for that, it is normally far lower than the cost of replacing the enterprise systems built around the model.

If that is roughly right, then a strategy built on keeping weights closed and centralising services behind proprietary infrastructure is defending the wrong thing. It protects a layer with a thinner moat than its price implies, while forfeiting the ecosystem effects that come from permissionless adoption. One nuance worth adding to the definition above: open weight does not automatically mean open source, unrestricted, or self hosted. It means those parameters are published under a licence that permits some degree of independent deployment and modification, which is narrower than it sounds, though still enough to let a model be hosted anywhere, altered, fine tuned, and embedded into workflows nobody at the vendor anticipated. That is close to the dynamic that made the open internet more valuable than any single walled service, and something like it is playing out again in AI. The problem with the moat thesis alone, and it matters for everything that follows, is that a thin commercial moat is only half the story. Layer an involuntary regulatory one on top of it, and the case for staying closed and centralised gets worse, not better.

2. The evidence is compounding

Werdmuller’s post links to Robert Hart’s reporting in The Verge on how tight the gap has become between Chinese open models and the American frontier. The Kimi K3 launch is a live data point for that argument, though it needs one correction before it can carry any weight. Moonshot launched Kimi K3 on 16 July, a 2.8 trillion parameter model, and within 48 hours had to pause new signups because demand pushed its GPUs to the limit. That is evidence of extraordinary demand for Chinese frontier capability. It is not yet proof of open weight adoption, because Moonshot is currently serving K3 only through its own hosted API and has committed to releasing the full weights on 27 July. What the pause shows is that the market has stopped treating Chinese models as second tier alternatives, worth watching again once the weights are actually out and independently run.

The broader numbers back this up, with the same care about what they actually measure. One analysis cited by Lawfare estimates that on OpenRouter, a large model routing platform, Chinese open weight models grew from under 2 percent of token traffic in late 2024 to roughly 61 percent of usage within its group of top models by mid 2026. That is not a measure of the whole AI market or of enterprise deployment generally, but it is a real signal of developer adoption. A16z partner Martin Casado has separately put the odds that any given startup is using a Chinese model somewhere in its stack at around 80 percent, which says something about presence, not necessarily about which model carries the workload or the revenue. Neither number proves the market has been won. Both point the same direction.

3. The case the other side is making

To be fair to the US position, none of this is happening because Washington woke up one day and decided to make life difficult for its own AI industry. The Commerce Department directive that pulled Fable 5 and Mythos 5 offline was triggered by a reported jailbreak, with Amazon researchers flagging that the model’s cyber capabilities could be elicited in ways the guardrails were meant to prevent. The controls were lifted on 30 June and Fable 5 was back globally on 1 July, so the outage itself lasted nineteen days, not indefinitely. That limits the scale of the disruption, but not its significance: customers discovered that access to a production model could disappear worldwide, immediately, over a decision neither they nor the vendor controlled. That is a real concern, not a manufactured one, and it sits alongside a genuinely difficult problem nobody has solved yet: at frontier capability, a model that is good enough to help a defender is also good enough to help an attacker, and no one has built a reliable way to tell the two apart at the point of the API call.

The same tension shows up in reverse a few months earlier, though it is worth being precise about which guardrails were in dispute each time. Anthropic was designated a supply chain risk by the Pentagon after refusing to remove restrictions on its models being used for mass domestic surveillance and fully autonomous weapons, and lost federal contracts to OpenAI, Google, and Microsoft as a result. That was a dispute about permitted use cases. The Fable 5 suspension a few months later was a dispute about a different set of guardrails entirely, the cybersecurity safeguards meant to stop the model’s capabilities being extracted through a jailbreak. These are not the same guardrails, but within a few months the same company was penalised from opposite directions: once for refusing to loosen restrictions the government wanted gone, and once because a different set of restrictions reportedly did not hold. That is not a coherent policy. But it is evidence that the underlying question, how much capability to allow, to whom, and under what oversight, is a real one that the open weight side of this argument does not get to skip just because it is inconvenient.

4. Regulatory strangulation, not just competition

Where I think the US position falls apart is not the underlying safety concern. It is the incoherence of the mechanism. The administration had already rejected the Biden era AI Diffusion Rule, the licensing regime that governed export of advanced chips and closed model weights, back in May 2025, on the grounds that it was overly bureaucratic and would stifle American competitiveness. Just over a year later, in June 2026, it reached for that same export control authority, applied for the first time to a live software inference endpoint rather than a physical good, to pull two commercially deployed models offline worldwide with almost no notice. Those two actions are not formally contradictory, since they happened under different circumstances a year apart. But together they show how quickly access to a model can become subject to discretionary government intervention, in either direction, regardless of what the licensing rulebook said the week before.

It has not stayed contained to Anthropic, and the strongest evidence for that is OpenAI’s own account of what happened next. OpenAI limited the rollout of GPT-5.6 to roughly 20 government approved partners at the administration’s request, ahead of a broader release. The company complied, and then said plainly, in its own announcement, that it does not believe this kind of government access process should become the long term default, warning that it keeps capable tools away from the developers, enterprises, and cyber defenders who need them. That is a frontier lab on record saying the mechanism itself is the problem, not just critics of the industry.

Google’s case is murkier and I want to be careful not to overstate it. Gemini 3.5 Pro missed the June release window the company announced at its developer conference in May, and subsequent reporting suggests the model underwent substantial additional training and has missed further internal targets since, though Google has not confirmed the specifics publicly. What Google has confirmed is that it is discussing the model’s capabilities and testing standards with the US government ahead of release. Whatever the underlying cause of the delay, a major lab’s launch timeline now running through a conversation with a federal agency, on top of its own technical problems, is the more defensible point here. The mechanism has also started to strain relationships beyond China. Lawmakers have introduced a bill extending US export restrictions to Dutch and Japanese lithography equipment, which allied governments have described as a legislative mandate imposed on them rather than something negotiated with them. A policy built to contain a strategic rival is now creating friction with the allies that rival is supposed to be isolated from.

Contrast all of that with the posture from the second screenshot. That team’s forensic work needed a model that would not treat every legitimate request as a threat, and needed infrastructure that did not send attacker artefacts or referenced credentials outside the organisation’s own boundary. Being open weight was necessary but not sufficient here. What actually gave them both of those things was self hosting the model inside infrastructure they controlled, rather than routing sensitive material through anyone else’s endpoint, Chinese, American, or otherwise. No vendor guardrail in the loop deciding whether an incident responder looked too much like an attacker. Once the model runs inside the organisation’s own boundary, no incident data needs to leave that boundary for a third party to inspect, log, or restrict. That is not a hypothetical enterprise risk scenario. It is a team that tried the commercial route first, hit a wall built by safety tooling that cannot distinguish defenders from attackers, and solved the problem by taking control of the stack.

5. What the market movement does and does not tell us

Chipmakers and other AI infrastructure names have slumped repeatedly through mid and late July, with the semiconductor index down close to 10 percent over the month. That is driven by a mix of ordinary factors, profit taking after a long run up, unease about capital expenditure now projected to top a trillion dollars across the industry by 2027, and shifting interest rate expectations, and open weight competition does not explain it on its own. What it does do is pressure the pricing and margin assumptions built into proprietary inference economics. If capable models are becoming substitutable at the same time government intervention makes access to any single model less predictable, the return profile behind current valuations gets harder to defend, on top of the more conventional explanations rather than instead of them.

6. What this means if you are buying, not building

I keep coming back to the question I raised in the jurisdiction piece: not whether a company can physically move its headquarters out of American jurisdiction, but whether the regulatory trajectory makes the jurisdiction structurally hostile to the kind of company you need it to be. The Kimi K3 and GLM 5.2 examples suggest the same question now applies one layer up, to the enterprises buying AI rather than building it.

If your vendor’s model can be pulled from under you by a Commerce Department letter with a few hours notice, that is a concentration risk you did not choose to accept. If your vendor’s safety tooling cannot distinguish your own security team from an attacker, that is an operational risk with a real cost attached, measured in the hours your incident responders spend arguing with a guardrail instead of doing the analysis. Open weight infrastructure that you host and control does not eliminate governance and safety obligations, and the concerns raised in section 3 do not go away just because you self host. It moves those obligations to where you can actually manage them, rather than leaving them subject to a regulator or a vendor’s risk appetite that has nothing to do with your own.

I don’t think the right conclusion is ideological loyalty to open weight over closed models. It is substitutability. Keep the model behind an abstraction layer rather than hard coded to one vendor’s API. Run workload specific evaluations across more than one provider, so switching is a tested option rather than a theoretical one. Preserve a self hosted path for the sensitive workloads where a vendor’s refusal behaviour or a government directive would otherwise be a single point of failure. Track legitimate refusal rates alongside benchmark quality, the way that team tracked how often their own security work got blocked. The strategic value of open weight models is not that every enterprise needs to run a 2.8 trillion parameter model on its own hardware. It is that their existence gives you somewhere else to go, which is the one thing a closed, centralised, single vendor stack structurally cannot offer you.

That is the thread connecting a Chinese subscription pause, a security team’s forensic workaround, and a slump in American AI infrastructure stocks. Closed and centralised was always a bet that the moat existed and that access would stay predictable. Increasingly, neither assumption looks safe to make on its own.