The AI race has split in two
Two closed flagships shipped in three days and split the frontier — Claude Fable 5.1 leads Humanity's Last Exam at 59.1%, GPT-6 Astra leads Terminal-Bench 4.0 at 58.2%. Kimi K3 holds the open line 12.2 points back, open weights carry ~69% of routed tokens, and the August open-weight promises finally shipped — under bespoke licences.
Best closed model
Claude Fable 5.1 · HLE, no tools — set Sep 1; GPT-6 Astra landed 4.4 pts below it two days later
Best open model
Kimi K3 · HLE — two months old and still unbeaten by any open release
Open–closed frontier gap
Two September flagships moved the closed line; the open line moved with a board re-run, not a release
Open share of routed tokens
Of ~115T tokens a week on OpenRouter; Chinese-developed models hold six of the top ten slots
Frontier capability over time
Best available score on Humanity's Last Exam (no tools) — closed vs open weights. Interior points are approximate (≈); endpoints verified September 15, 2026.
How far behind is open, benchmark by benchmark?
Best generally available closed vs best open score. Shorter connector = smaller gap.
The price of the frontier
SWE-bench Pro score vs published output price per 1M tokens (log scale). Models without a published price or score are excluded, not estimated.
OpenRouter: open vs closed token share
Share of routed tokens. Endpoints reported; interior points approximate (≈).
Hugging Face downloads by lab origin
Share of open-model downloads by publishing lab's region. Endpoints from Hugging Face reporting; interior points approximate (≈).
The last 30 days
Every chart above moved because of these eight events.
Alibaba publishes the first Max-tier Qwen weights
Qwen3.8-2.4T-A95B goes up nine days after launch under a bespoke licence with a $50M revenue gate — text-only, without the API's vision or 1M window.
DeepSeek V4-Pro leaves preview; Google ships Gemini 3.7 Flash
V4-Pro GA lands as an agent release with peak-hour output pricing up to $3.96/1M, against a flat $0.87. Google's second Flash release in three weeks.
Z.ai launches GLM-5.3
Same 753B/40B base as 5.2, ~50% better coding, CyberGym 77.2 → 84.5. Weights follow on Aug 28.
GLM-5.3 weights published — under a review clause
The flagship drops MIT for a bespoke licence requiring Z.ai security review above $10B revenue. MIT now lives at the Flash tier, where GLM-5.3-Flash is #2 on OpenRouter.
Anthropic ships Claude Fable 5.1 and Mythos 5.1
59.1% HLE and 81.2% SWE-bench Pro at Fable 5's list price, with cache reads cut 75% — ~25% cheaper on typical workloads, ~45% on agentic ones.
Gemini 3.8 Flash and Muse Spark 1.3 land the same day
Google's third Flash in six weeks holds price with a printed expiry; Meta's coding flagship adds a contributor endpoint at ~$0.10/$0.20 if you let Meta train on your traffic.
GPT-6 Astra ships; NVIDIA agrees to buy Hugging Face
Astra goes to Daybreak trusted-access first, GA Sep 8, at $10/$50 — cyber capability gated at launch. Hours later NVIDIA confirms a ~$12.9B deal for the Hub.
Terminal-Bench 4.0 snapshot resets the terminal ladder
Eighteen models, five trials per task, published variance — and an 87.6 → 19.1 collapse for Gemini 3.8 Flash. GLM-5.3 is the top open entry at 41.8%.
The story in four beats
The frontier changed hands twice in three days
Claude Fable 5.1 shipped September 1 and GPT-6 Astra September 3, and they split the scoreboard: Astra leads Terminal-Bench 4.0 by 0.3 points, Fable 5.1 leads HLE by 4.4 and SWE-bench Pro outright. There is no single frontier model this cycle — there is a terminal-agent leader and a reasoning leader, and they are different companies' flagships at the same $10/$50 list price.
Terminal-Bench 4.0 deleted the old coding ladder
Gemini 3.8 Flash scored 87.6% on Terminal-Bench 2.1 and 19.1% on 4.0. Grok 4.5 went from 83.3% to 12.4%. The new harness runs five independent trials per task and publishes variance, and the top score fell from ~90% to 58.2%. Anyone who selected a model on 2.1 numbers in the last two months selected on a benchmark the field had learned.
The open-weight promises landed — under bespoke licences
Alibaba published Qwen3.8-Max's weights on August 12 and Z.ai published GLM-5.3's on August 28, so the two biggest August commitments are real downloads. Both carry custom licences with revenue gates, the Qwen checkpoint is text-only, and MIT has retreated to the Flash tier. Meta's Muse Spark 1.2 weights, promised August 10 'in the coming weeks', still do not exist.
NVIDIA bought the open-model distribution layer
On September 3, NVIDIA agreed to acquire Hugging Face for about $12.9 billion — 18 million developers and 3 million models — while OpenRouter ran ~115 trillion tokens in a single week with open weights at 69% of them. The infrastructure the open ecosystem depends on is consolidating at the same rate the models themselves are commoditizing.
Go deeper
Model roster
Claude Fable 5.1, GPT-6 Astra, Kimi K3, and the rest — context, pricing, modalities, and license nuance.
ExploreBenchmarks
Humanity's Last Exam, agentic coding, and terminal agents — plus which benchmarks are already saturated.
ExploreEcosystem
Where tokens get routed on OpenRouter and which labs the world is downloading on Hugging Face.
Explore