Ecosystem signals
Capability is one thing; adoption is another. Two lenses on demand — where tokens get routed in production (OpenRouter) and which model weights the world is pulling down (Hugging Face). Verified 2026-09-15.
OpenRouter — routed usage
115T
tokens routed in one week
Week of Aug 31–Sep 6, against ~25T/week in May — the router grew 4.6× in four months.
69.1%
of routed tokens go to open weights
Against 30.9% closed, measured within the daily top 50 by the IAPS-AI dashboard. Roughly flat since July.
6 of 10
top models by volume are Chinese-developed
Down from eight in August — because GPT-5.6 Luna jumped 129% week over week into #3, not because Chinese volume fell.
~12%
Anthropic token share — but the revenue lead
Premium per-token pricing means Claude captures the most dollars while open models take the volume.
Open vs closed token share
Share of routed tokens. Endpoints reported; interior points approximate (≈).
Volume is not value
Why both sides claim they're winning.
Token share measures volume, not value. Cheap open models absorb bulk workloads (extraction, summarization, high-volume agents) while premium closed models keep the tasks where a bad output costs real money. Both statements are true at once: open weights run about 69% of routed tokens; closed models earn most of the revenue. What is new this month is the shape of the competition at the bottom. GLM-5.3-Flash went from launch to ~10T tokens a week — the fastest climb on the router this year — while OpenAI's GPT-5.6 Luna rose 129% in a week to #3 at $0.20/$1.20, the first closed model to compete on open-weight terms. Meanwhile the cheap end started charging like the expensive end: DeepSeek's V4-Pro GA introduced peak-hour output pricing up to $3.96 per 1M against a flat $0.87, and Meta priced Muse Spark 1.3's contributor tier in training data rather than dollars. The split is no longer open-versus-closed on price; it is who is willing to sell tokens below cost, and what they want in return.
The tell in the data
Anthropic holds roughly 12% of tokens and still captures the most revenue on the router, while the #1 model by volume — DeepSeek V4 Flash at ~12.1T tokens a week — serves at $0.28 per 1M output, about 0.6% of Claude Fable 5.1's list price. The new wrinkle is that cheap no longer means flat: DeepSeek's V4-Pro GA added peak-hour output pricing up to $3.96 per 1M, and Meta will sell you Muse Spark 1.3 at a tenth of list if you let it train on your traffic.
Hugging Face — open-model downloads
$12.9B
NVIDIA's agreed price for Hugging Face
Announced Sep 3, closing expected H1 2027. The largest publisher of open weights now also owns where the world downloads them.
41%
of past-year Hub downloads are Chinese models
Over 45% on recently uploaded models — Hugging Face reports China overtook the US on recent adoption.
3B+
cumulative Qwen downloads
Against 418M for Google and 227M for Meta in 2026 — and 113k+ derivative models, more than both combined.
2 of 2
August open-weight flagships carry bespoke licences
Qwen3.8-Max (Aug 12) and GLM-5.3 (Aug 28) both shipped weights with revenue gates rather than Apache-2.0 or MIT.
36 days
since Meta promised Muse Spark 1.2's weights
Announced Aug 10 'in the coming weeks'. No repository, no parameter count, no licence as of this pass.
Download share by lab origin
Share of open-model downloads by publishing lab's region. Endpoints from Hugging Face reporting; interior points approximate (≈).
Trending on the Hub right now
After August's two flagship weight drops — Qwen3.8-Max and GLM-5.3, both under bespoke licences. Download counts move daily, so ranks are shown without stale precision.
GLM-5.3-Flash
MIT; 320B/18B and #2 on OpenRouter by volume
GLM-5.3
Weights Aug 28; bespoke licence with a security-review clause
Qwen3.8-2.4T-A95B
First Max-tier Qwen weights; text-only checkpoint
Kimi K3
2.8T weights; still the best open model on HLE
Muse Glimmer 30B
Apache-2.0; one consumer GPU
Nemotron 3 Ultra
Strongest US-origin open model on the router
4 of the 6 trending models are from Chinese labs, and 3 of the 4 carry bespoke licences with revenue gates. The permissive rows are now GLM-5.3-Flash (MIT) and Muse Glimmer 30B (Apache-2.0) — a Flash tier and a 30B local model. No current flagship ships under an OSI licence.