The AI race has split in two
Closed models still hold the frontier — Claude Opus 5 leads Humanity's Last Exam at 56.3% — but Kimi K3 cut the open gap to 12.8 points, the narrowest on record, while open weights carry ~71% of production tokens. Capability and adoption are converging faster than either side planned.
Best closed model
Claude Opus 5 · HLE, no tools — shipped Jul 24 at half Fable 5's price
Best open model
Kimi K3 · HLE — 2.8T open weights, the largest model released to date
Open–closed frontier gap
Narrowest reading on record — Kimi K3 closed it faster than Opus 5 reopened it
Open share of routed tokens
Chinese-origin models alone are past 60% and now hold all five top OpenRouter slots
Frontier capability over time
Best available score on Humanity's Last Exam (no tools) — closed vs open weights. Interior points are approximate (≈); endpoints verified July 30, 2026.
How far behind is open, benchmark by benchmark?
Best generally available closed vs best open score. Shorter connector = smaller gap.
The price of the frontier
SWE-bench Verified score vs published output price per 1M tokens (log scale). Models without a published price or score are excluded, not estimated.
OpenRouter: open vs closed token share
Share of routed tokens. Endpoints reported; interior points approximate (≈).
Hugging Face downloads by lab origin
Share of open-model downloads by publishing lab's region. Endpoints from Hugging Face reporting; interior points approximate (≈).
The last 30 days
Every chart above moved because of these eight events.
Anthropic ships Claude Sonnet 5
80.4% on Terminal-Bench 2.1 — past Opus 4.8 — at $2/$10 introductory pricing; default for Free and Pro.
Beijing weighs export controls on Chinese models
Reuters reports meetings with Alibaba, ByteDance, and Z.ai on restricting overseas access — open-weight releases included.
xAI ships Grok 4.5
#1 on agentic tool use and #4 overall on the AA index; 83.3% Terminal-Bench 2.1 at $2/$6 and a 500K window.
GPT-5.6 Sol, Terra, and Luna go GA
Sol takes Terminal-Bench 2.1 SOTA at 88.8%; its HLE (47.2%) and SWE-bench (96.2%) scores land later in the month.
Moonshot releases Kimi K3
2.8T open weights, 1M context — 43.5% HLE and 93.4% SWE-bench Verified, and first place on Frontend Code Arena.
Google ships Gemini 3.6 Flash
Scores 50 on the AA index — above Google's own 3.1 Pro Preview — at $1.50/$7.50 using ~17% fewer output tokens.
Claude Opus 5 ships; open-weight coalition letter lands
Opus 5 takes HLE (56.3%) and Terminal-Bench 2.1 (89.1%) at Opus 4.8's price. The same day, 25 companies led by Nvidia and Microsoft publish an open-weight advocacy letter.
Chinese models sweep the OpenRouter top five
A first for the router, led by Xiaomi's MiMo-V2.5 — capability leadership and volume leadership are now on different continents.
The story in four beats
The open gap shrank by a quarter in eight days
Kimi K3 shipped July 16 with 2.8 trillion open weights and took the open HLE record from ≈36% to 43.5%, plus 93.4% on SWE-bench Verified. Opus 5 answered on July 24 with 56.3%. Net: the gap fell from 17.3 points to 12.8 — the narrowest on record, and it narrowed while the frontier was still moving.
Frontier capability got cheaper, not dearer
Claude Opus 5 beats Fable 5 on every headline benchmark at $5/$25 — half the price. GPT-5.6 Terra matches GPT-5.5 at half. Gemini 3.6 Flash beats Google's own Pro preview on the AA index while cutting its output price. Three labs, one month, the same move.
Chinese models swept the router
For the first time, Chinese-developed models hold all five top OpenRouter slots by volume, led by Xiaomi's MiMo-V2.5 — a 15B-active model serving at roughly $0.22 per 1M output tokens. Chinese-origin models now carry over 60% of routed traffic on a router doing more than 20 trillion tokens a week.
Both governments reached for the same lever
Washington took Fable 5 offline for 19 days in June; Beijing met with Alibaba, ByteDance, and Z.ai on July 7 about restricting overseas access to Chinese frontier models — open weights included. A 25-company Nvidia/Microsoft coalition pushed back publicly on July 24. Openness is now a policy variable on both sides of the Pacific.
Go deeper
Model roster
Claude Opus 5, GPT-5.6 Sol, Kimi K3, and the rest — context, pricing, modalities, and license nuance.
ExploreBenchmarks
Humanity's Last Exam, agentic coding, and terminal agents — plus which benchmarks are already saturated.
ExploreEcosystem
Where tokens get routed on OpenRouter and which labs the world is downloading on Hugging Face.
Explore