Verified snapshot · July 9, 2026

The AI race has split in two

Closed models hold the frontier — Claude Fable 5 leads Humanity's Last Exam at 53.3% — while open weights now carry 69% of production tokens at a fraction of the price. Capability and adoption are no longer the same race.

Best closed model

53.3%9 pts vs Q1

Claude Fable 5 · HLE, no tools — redeployed Jul 1 after a 19-day export-control pause

Best open model

≈36%9 pts vs Q1

GLM-5.2 (MIT) · HLE — Artificial Analysis puts the top open cluster at 34–36%

Open–closed frontier gap

17.3 pts0.4 pts vs Q1

Fell to ~9 pts while Fable 5 was offline in June; snapped back on Jul 1

Open share of routed tokens

69.1%47 pts YoY ≈

Open passed closed on OpenRouter — DeepSeek alone routes 16.3%, more than any provider

Frontier capability over time

Best available score on Humanity's Last Exam (no tools) — closed vs open weights. Interior points are approximate (≈); endpoints verified July 9, 2026.

Closed / proprietary
Open weights & source
The 2026 Q2 closed point is Gemini 3.1 Pro — Fable 5 posted 53.3% on June 9 but spent June 12–30 offline under a US export-control order. The gap read ~9 points while it was dark and snapped to 17.3 the day it returned: the frontier now moves by release (and by regulator), not by trend line.

How far behind is open, benchmark by benchmark?

Best generally available closed vs best open score. Shorter connector = smaller gap.

Closed / proprietary
Open weights & source
The gap tracks how open-ended the work is: 4.3 pts on graduate science (saturating), 14.4 on agentic coding, 17.3 on frontier reasoning. Restricted-access Claude Mythos 5 (95.5% SWE-bench) would stretch the coding gap further.
vals.ai · llm-stats · Artificial Analysis, Jul 2026

The price of the frontier

SWE-bench Verified score vs published output price per 1M tokens (log scale). Models without a published price or score are excluded, not estimated.

Closed / proprietary
Open weights & source
DeepSeek V4 Pro Max ties Gemini-tier coding at $0.87 per 1M output tokens — under 2% of Fable 5's list price. Above 88% the chart belongs to Anthropic; below it, the price-performance frontier is entirely open weights. (GPT-5.6 Sol went GA today at $30 output; its SWE-bench score isn't published yet.)
Provider list prices · vals.ai / BenchLM scores, Jul 2026

OpenRouter: open vs closed token share

Share of routed tokens. Endpoints reported; interior points approximate (≈).

Closed / proprietary
Open weights & source
Open weights crossed 50% around Q1 2026 and now run 69.1% of routed tokens — a flip driven by price, not leaderboards. It happened without open models holding a single #1 benchmark slot.

Hugging Face downloads by lab origin

Share of open-model downloads by publishing lab's region. Endpoints from Hugging Face reporting; interior points approximate (≈).

China
US
Europe
Other
Hugging Face reports Chinese labs at 41% of past-year downloads and over 45% on recently uploaded models — the overtake already happened on new releases. Qwen alone passed 1B cumulative downloads in March.

The last 30 days

Every chart above moved because of these six events.

Jun 9

Anthropic ships Claude Fable 5 and Mythos 5

First Mythos-class models: 95.0% SWE-bench Verified (independently confirmed), new HLE high.

Jun 12

US export controls take Fable 5 and Mythos 5 offline

Order followed an Amazon researchers' jailbreak report; access suspended globally for 19 days.

Jun 13

Z.ai releases GLM-5.2 under MIT

First open-weight model to beat GPT-5.5 on SWE-bench Pro (62.1%); new open leader on HLE and the AA index.

Jun 26

OpenAI previews GPT-5.6; Mythos 5 partially restored

Sol/Terra/Luna preview begins; ~100 US critical-infrastructure orgs regain Mythos 5 access.

Jun 30

Commerce lifts the export-control order

Anthropic redeploys Fable 5 globally on Jul 1 with a new cybersecurity classifier; it retakes #1 on HLE.

Jul 9

GPT-5.6 Sol, Terra, and Luna go GA

Sol at $5/$30 with Terminal-Bench 2.1 SOTA (88.8%); HLE and SWE-bench scores not yet published.

Closed-model release
Open-model release
Policy

The story in four beats

01

The frontier answers to regulators now

Claude Fable 5 shipped June 9, was taken offline June 12 by a US export-control order tied to a jailbreak disclosure, and returned July 1 — the first time the leading model spent 19 days dark by government order. Frontier capability is now regulated infrastructure.

02

GPT-5.6 landed today

Sol went GA July 9 at $5/$30 per 1M tokens with a new max reasoning effort and a subagent 'ultra mode', setting state of the art on Terminal-Bench 2.1 (88.8%). Terra matches GPT-5.5 at roughly half the price — the frontier price war is now explicit.

03

Open weights won the volume war

69.1% of OpenRouter tokens now route to open models; Chinese open weights hold six of the top ten slots. Anthropic runs ~12% of tokens yet captures the most revenue — volume and value have split.

04

The gap is release-driven, not secular

The HLE gap compressed to ~9 points in June (GLM-5.2 vs Gemini 3.1 Pro), then snapped to 17.3 when Fable 5 returned. Expect it to compress again on the next open flagship — the trend line moves in release-sized steps.

Go deeper