Illustration: AI-generated · Inside China AI
Open Weights, Two Meanings: What China’s Four-Model Month Reveals
This analysis is also available in German: Zur deutschen Fassung
Between late July and late August 2026, four Chinese laboratories put frontier-scale open-weight models on Hugging Face. DeepSeek released V4-Flash on 31 July. Alibaba followed with Qwen3.8 in the second week of August. Zhipu shipped GLM-5.3-Flash on 25 August. Moonshot’s Kimi K3 sits alongside them. What had been a quarterly rhythm compressed into roughly a month.
The speed is the obvious story. It is not the interesting one. The interesting one is that „open weights“ has quietly split into two economically different things — and the download figures suggest the market has already noticed.
1. Four models, four different bets
These are not iterations of a common design. Each lab made a distinct architectural wager, and the parameter counts published on Hugging Face show how far apart they are:
- Kimi K3 (Moonshot) — 2.78 trillion parameters, the largest of the group. Its model card describes a Stable LatentMoE framework that „activates 16 out of 896 experts“, the widest expert pool among the four, alongside a Delta Attention mechanism (model card).
- Qwen3.8-Max (Alibaba) — 2.4 trillion parameters with 95 billion active, a natively multimodal mixture-of-experts model with a one-million-token context window, announced on 3 August with open weights following roughly ten days later (announced by Alibaba and widely reported in trade coverage; we could not verify the figures against the repository itself, because the weights are licence-gated and return an authorisation error without prior acceptance of terms).
- DeepSeek V4-Flash — 304 billion parameters, by far the smallest, and explicitly built for inference speed. The card documents DSpark speculative decoding with a draft module shipped inside the checkpoint (model card).
- GLM-5.3-Flash (Zhipu) — 321 billion parameters, the model that briefly ran on OpenRouter under the anonymous name Ox Alpha before Zhipu claimed it (model card). We covered that episode in Ox Alpha Was Chinese.
The spread is the point: from 304 billion to 2.78 trillion parameters, from maximal scale to maximal inference economy. Four labs, four readings of what „frontier“ should mean.
2. The split that matters: MIT against revenue-gated
Licensing divides the group cleanly in two, and the division does not follow model size.
DeepSeek V4-Flash and GLM-5.3-Flash are both released under the MIT licence. That is the most permissive of the mainstream open-source licences: download, modify, self-host, deploy commercially, no negotiation, no threshold, no notification.
Kimi K3 and Qwen3.8-Max carry custom licences that gate commercial use. Moonshot’s licence text is explicit about where the gate sits: if a licensee operates a model-as-a-service business and aggregate revenue „exceeds 20 million US dollars … in total over any consecutive 12 months“, the licensee „must enter into a separate agreement with Moonshot“ (licence text). Alibaba’s flagship repository is gated: it cannot be read without first accepting terms.
This is not open source in the sense the term normally carries. It is open distribution with a deferred commercial claim — free until you succeed, then chargeable. For a startup, that is an attractive proposition. For a company that scales, it is a liability that materialises at precisely the moment migration is hardest.
3. What the download figures reveal
Adoption data is the part of this story that requires no interpretation. These are Hugging Face’s own thirty-day download counts, retrieved on 1 September 2026:
- DeepSeek V4-Flash (MIT): 4,650,353
- Qwen3.8-27B (Apache 2.0): 4,960,483
- Kimi K3 (custom licence): 2,783,061
- GLM-5.3-Flash (MIT): 441,348
The sharpest comparison sits inside a single laboratory. Alibaba released two models within days of each other: the 2.4-trillion-parameter flagship under a custom, gated licence, and a 27-billion-parameter model under Apache 2.0. The small permissive one has been downloaded almost five million times. The flagship cannot be downloaded at all without accepting terms first.
Same lab, same month, same reputation behind both. The variable that differs is the licence — and the gap in adoption is two orders of magnitude.
A caveat on the figures: Hugging Face counts downloads, not deployments, and a gated repository is structurally disadvantaged in any such count. The numbers describe reach and friction, not revenue or production use. But friction is exactly what a licence is designed to create, and here it is measurable.
4. What the benchmarks do and do not say
Moonshot’s model card places Kimi K3 at 88.3 on Terminal Bench 2.1 and 93.5 on GPQA Diamond, with 67.5 on DeepSWE, a coding-agent benchmark designed to resist data contamination, against 54.4 for DeepSeek V4-Flash.
Two qualifications belong next to those numbers, and they are not decoration. First, they are vendor-reported. Second — and this is the part usually omitted — the comparison table is Moonshot’s own measurement of its competitors, not each laboratory’s published figure for itself. A vendor’s scoring of rival models is evidence of what that vendor wants shown, not an independent ranking.
What can be said without a leaderboard: on contamination-resistant agentic benchmarks the spread between these four models is narrower than on classical knowledge tests. The interesting competition has moved from what a model knows to what it can carry out.
Why it matters for Europe
Two of these models are, for European purposes, free infrastructure. A frontier-class model under MIT or Apache 2.0 can be downloaded, modified, self-hosted and commercially deployed by a European hospital, public administration, bank or defence supplier without negotiating with anyone. For organisations bound by data-residency rules that rule out foreign APIs, that is not a marginal option — it is the difference between having a frontier model on their own infrastructure and not having one.
The other two carry a bill that arrives late. A revenue-gated licence is invisible during the pilot, invisible during the rollout, and becomes decisive at the moment a company scales — which is also the moment its data, prompts, fine-tunes and internal tooling are hardest to move. European procurement rarely asks about licence thresholds during a proof of concept. On this evidence, it should. The relevant question is not „what does the licence permit today“ but „what does it require at ten times our current volume“.
And there is a strategic reading. Chinese laboratories are running both models of openness at once: permissive licences to build ecosystem reach, gated licences to convert that reach into revenue. Whichever wins, the experiment is being conducted at frontier scale and in public — which is more than can be said for the closed American systems these models are measured against. Europe is currently a consumer of that experiment rather than a participant in it.
All parameter counts, licences and download figures were retrieved from the Hugging Face API on 1 September 2026 and are linked to their model cards. Benchmark figures are vendor-reported and identified as such. Corrections and additional information are welcome: press@insidechinaai.com · Our editorial standards.


No responses yet