Verified snapshot · September 15, 2026

The AI race has split in two

Two closed flagships shipped in three days and split the frontier — Claude Fable 5.1 leads Humanity's Last Exam at 59.1%, GPT-6 Astra leads Terminal-Bench 4.0 at 58.2%. Kimi K3 holds the open line 12.2 points back, open weights carry ~69% of routed tokens, and the August open-weight promises finally shipped — under bespoke licences.

Best closed model

59.1%2.8 pts vs Aug 13

Claude Fable 5.1 · HLE, no tools — set Sep 1; GPT-6 Astra landed 4.4 pts below it two days later

Best open model

46.9%3.4 pts vs Aug 13

Kimi K3 · HLE — two months old and still unbeaten by any open release

Open–closed frontier gap

12.2 pts0.6 pts vs Aug 13

Two September flagships moved the closed line; the open line moved with a board re-run, not a release

Open share of routed tokens

69%28 pts YoY ≈

Of ~115T tokens a week on OpenRouter; Chinese-developed models hold six of the top ten slots

Frontier capability over time

Best available score on Humanity's Last Exam (no tools) — closed vs open weights. Interior points are approximate (≈); endpoints verified September 15, 2026.

Closed / proprietary
Open weights & source
Both lines moved this cycle, but only one moved because of a model. Claude Fable 5.1 took the closed line to 59.1% on September 1; the open line rose to 46.9% when the board re-ran Kimi K3, which has not been passed by an open release since July. Read the last step with that caveat — the same re-run moved Claude Opus 5 down 1.4 points without Anthropic touching it.

How far behind is open, benchmark by benchmark?

Best generally available closed vs best open score. Shorter connector = smaller gap.

Closed / proprietary
Open weights & source
Two of these three ladders are new, and the harder harnesses reopened gaps that the saturated ones had hidden: 13.5 points on agentic coding (SWE-bench Pro, not Verified) and 16.4 on terminal agents (Terminal-Bench 4.0, not 2.1). Frontier reasoning is the narrowest at 12.2. The open rows are Qwen3.8-Max, GLM-5.3, and Kimi K3 — and for the first two, August turned a licence promise into an actual download.
BenchLM · tbench.ai · pricepertoken, Sep 2026

The price of the frontier

SWE-bench Pro score vs published output price per 1M tokens (log scale). Models without a published price or score are excluded, not estimated.

Closed / proprietary
Open weights & source
On the harder coding harness the frontier premium reappears: Fable 5.1 resolves 81.2% at $50 per 1M output, and the best open row — Qwen3.8-Max at 67.7% — sits 13.5 points back at 12% of the price. That is a wider capability gap than SWE-bench Verified showed all summer, and a much wider price gap, because Verified had compressed everything above 93% into a single band. The open point that matters is that Qwen3.8-Max's weights actually published in August; in the July chart the same capability band was occupied by models you could only rent.
Provider list prices · BenchLM SWE-bench Pro scores, Sep 2026

OpenRouter: open vs closed token share

Share of routed tokens. Endpoints reported; interior points approximate (≈).

Closed / proprietary
Open weights & source
Open weights crossed 50% around Q1 2026 and have held near 69% since July, on a router that ran ~115T tokens in the week of Aug 31–Sep 6. Chinese-developed models hold six of the top ten slots by volume, down from eight — because OpenAI's GPT-5.6 Luna jumped 129% week over week into #3 at $0.20/$1.20, not because Chinese volume fell.

Hugging Face downloads by lab origin

Share of open-model downloads by publishing lab's region. Endpoints from Hugging Face reporting; interior points approximate (≈).

China
US
Europe
Other
Hugging Face reports Chinese labs at 41% of past-year downloads and over 45% on recently uploaded models — the overtake already happened on new releases. Qwen alone has passed 3B cumulative downloads against 418M for Google. On September 3, NVIDIA agreed to buy the Hub itself for ~$12.9B.

The last 30 days

Every chart above moved because of these eight events.

Aug 12

Alibaba publishes the first Max-tier Qwen weights

Qwen3.8-2.4T-A95B goes up nine days after launch under a bespoke licence with a $50M revenue gate — text-only, without the API's vision or 1M window.

Aug 13

DeepSeek V4-Pro leaves preview; Google ships Gemini 3.7 Flash

V4-Pro GA lands as an agent release with peak-hour output pricing up to $3.96/1M, against a flat $0.87. Google's second Flash release in three weeks.

Aug 14

Z.ai launches GLM-5.3

Same 753B/40B base as 5.2, ~50% better coding, CyberGym 77.2 → 84.5. Weights follow on Aug 28.

Aug 28

GLM-5.3 weights published — under a review clause

The flagship drops MIT for a bespoke licence requiring Z.ai security review above $10B revenue. MIT now lives at the Flash tier, where GLM-5.3-Flash is #2 on OpenRouter.

Sep 1

Anthropic ships Claude Fable 5.1 and Mythos 5.1

59.1% HLE and 81.2% SWE-bench Pro at Fable 5's list price, with cache reads cut 75% — ~25% cheaper on typical workloads, ~45% on agentic ones.

Sep 2

Gemini 3.8 Flash and Muse Spark 1.3 land the same day

Google's third Flash in six weeks holds price with a printed expiry; Meta's coding flagship adds a contributor endpoint at ~$0.10/$0.20 if you let Meta train on your traffic.

Sep 3

GPT-6 Astra ships; NVIDIA agrees to buy Hugging Face

Astra goes to Daybreak trusted-access first, GA Sep 8, at $10/$50 — cyber capability gated at launch. Hours later NVIDIA confirms a ~$12.9B deal for the Hub.

Sep 11

Terminal-Bench 4.0 snapshot resets the terminal ladder

Eighteen models, five trials per task, published variance — and an 87.6 → 19.1 collapse for Gemini 3.8 Flash. GLM-5.3 is the top open entry at 41.8%.

Closed-model release
Open-model release
Policy

The story in four beats

01

The frontier changed hands twice in three days

Claude Fable 5.1 shipped September 1 and GPT-6 Astra September 3, and they split the scoreboard: Astra leads Terminal-Bench 4.0 by 0.3 points, Fable 5.1 leads HLE by 4.4 and SWE-bench Pro outright. There is no single frontier model this cycle — there is a terminal-agent leader and a reasoning leader, and they are different companies' flagships at the same $10/$50 list price.

02

Terminal-Bench 4.0 deleted the old coding ladder

Gemini 3.8 Flash scored 87.6% on Terminal-Bench 2.1 and 19.1% on 4.0. Grok 4.5 went from 83.3% to 12.4%. The new harness runs five independent trials per task and publishes variance, and the top score fell from ~90% to 58.2%. Anyone who selected a model on 2.1 numbers in the last two months selected on a benchmark the field had learned.

03

The open-weight promises landed — under bespoke licences

Alibaba published Qwen3.8-Max's weights on August 12 and Z.ai published GLM-5.3's on August 28, so the two biggest August commitments are real downloads. Both carry custom licences with revenue gates, the Qwen checkpoint is text-only, and MIT has retreated to the Flash tier. Meta's Muse Spark 1.2 weights, promised August 10 'in the coming weeks', still do not exist.

04

NVIDIA bought the open-model distribution layer

On September 3, NVIDIA agreed to acquire Hugging Face for about $12.9 billion — 18 million developers and 3 million models — while OpenRouter ran ~115 trillion tokens in a single week with open weights at 69% of them. The infrastructure the open ecosystem depends on is consolidating at the same rate the models themselves are commoditizing.

Go deeper