Ecosystem signals

Capability is one thing; adoption is another. Two lenses on demand — where tokens get routed in production (OpenRouter) and which model weights the world is pulling down (Hugging Face). Verified 2026-09-15.

OpenRouter — routed usage

115T

tokens routed in one week

Week of Aug 31–Sep 6, against ~25T/week in May — the router grew 4.6× in four months.

69.1%

of routed tokens go to open weights

Against 30.9% closed, measured within the daily top 50 by the IAPS-AI dashboard. Roughly flat since July.

6 of 10

top models by volume are Chinese-developed

Down from eight in August — because GPT-5.6 Luna jumped 129% week over week into #3, not because Chinese volume fell.

~12%

Anthropic token share — but the revenue lead

Premium per-token pricing means Claude captures the most dollars while open models take the volume.

Open vs closed token share

Share of routed tokens. Endpoints reported; interior points approximate (≈).

Closed / proprietary
Open weights & source
Routed tokens are the strictest adoption test — real workloads, paid per token. Open's climb from ~22% to ~69% in a year happened without holding a single #1 leaderboard slot; price did the work. The line has been flat for two months, and the interesting movement is now inside it: GLM-5.3-Flash went from launch to ~10T tokens a week, while OpenAI's GPT-5.6 Luna became the first closed model to climb the router on open-weight pricing.

Volume is not value

Why both sides claim they're winning.

Token share measures volume, not value. Cheap open models absorb bulk workloads (extraction, summarization, high-volume agents) while premium closed models keep the tasks where a bad output costs real money. Both statements are true at once: open weights run about 69% of routed tokens; closed models earn most of the revenue. What is new this month is the shape of the competition at the bottom. GLM-5.3-Flash went from launch to ~10T tokens a week — the fastest climb on the router this year — while OpenAI's GPT-5.6 Luna rose 129% in a week to #3 at $0.20/$1.20, the first closed model to compete on open-weight terms. Meanwhile the cheap end started charging like the expensive end: DeepSeek's V4-Pro GA introduced peak-hour output pricing up to $3.96 per 1M against a flat $0.87, and Meta priced Muse Spark 1.3's contributor tier in training data rather than dollars. The split is no longer open-versus-closed on price; it is who is willing to sell tokens below cost, and what they want in return.

The tell in the data

Anthropic holds roughly 12% of tokens and still captures the most revenue on the router, while the #1 model by volume — DeepSeek V4 Flash at ~12.1T tokens a week — serves at $0.28 per 1M output, about 0.6% of Claude Fable 5.1's list price. The new wrinkle is that cheap no longer means flat: DeepSeek's V4-Pro GA added peak-hour output pricing up to $3.96 per 1M, and Meta will sell you Muse Spark 1.3 at a tenth of list if you let it train on your traffic.

Hugging Face — open-model downloads

$12.9B

NVIDIA's agreed price for Hugging Face

Announced Sep 3, closing expected H1 2027. The largest publisher of open weights now also owns where the world downloads them.

41%

of past-year Hub downloads are Chinese models

Over 45% on recently uploaded models — Hugging Face reports China overtook the US on recent adoption.

3B+

cumulative Qwen downloads

Against 418M for Google and 227M for Meta in 2026 — and 113k+ derivative models, more than both combined.

2 of 2

August open-weight flagships carry bespoke licences

Qwen3.8-Max (Aug 12) and GLM-5.3 (Aug 28) both shipped weights with revenue gates rather than Apache-2.0 or MIT.

36 days

since Meta promised Muse Spark 1.2's weights

Announced Aug 10 'in the coming weeks'. No repository, no parameter count, no licence as of this pass.

Download share by lab origin

Share of open-model downloads by publishing lab's region. Endpoints from Hugging Face reporting; interior points approximate (≈).

China
US
Europe
Other
Downloads measure who developers build on, not who ships the best demo. On recently uploaded models — the leading indicator — Chinese labs are already past 45%. As of September 3 the platform recording all of this is being acquired by NVIDIA for ~$12.9B, which makes the largest publisher of open weights the owner of the distribution channel too.

Trending on the Hub right now

After August's two flagship weight drops — Qwen3.8-Max and GLM-5.3, both under bespoke licences. Download counts move daily, so ranks are shown without stale precision.

1

GLM-5.3-Flash

MIT; 320B/18B and #2 on OpenRouter by volume

Z.ai
China
2

GLM-5.3

Weights Aug 28; bespoke licence with a security-review clause

Z.ai
China
3

Qwen3.8-2.4T-A95B

First Max-tier Qwen weights; text-only checkpoint

Alibaba
China
4

Kimi K3

2.8T weights; still the best open model on HLE

Moonshot
China
5

Muse Glimmer 30B

Apache-2.0; one consumer GPU

Meta
US
6

Nemotron 3 Ultra

Strongest US-origin open model on the router

NVIDIA
US

4 of the 6 trending models are from Chinese labs, and 3 of the 4 carry bespoke licences with revenue gates. The permissive rows are now GLM-5.3-Flash (MIT) and Muse Glimmer 30B (Apache-2.0) — a Flash tier and a 30B local model. No current flagship ships under an OSI licence.