The AI race has split in two
Closed models hold the frontier — Claude Fable 5 leads Humanity's Last Exam at 53.3% — while open weights now carry 69% of production tokens at a fraction of the price. Capability and adoption are no longer the same race.
Best closed model
Claude Fable 5 · HLE, no tools — redeployed Jul 1 after a 19-day export-control pause
Best open model
GLM-5.2 (MIT) · HLE — Artificial Analysis puts the top open cluster at 34–36%
Open–closed frontier gap
Fell to ~9 pts while Fable 5 was offline in June; snapped back on Jul 1
Open share of routed tokens
Open passed closed on OpenRouter — DeepSeek alone routes 16.3%, more than any provider
Frontier capability over time
Best available score on Humanity's Last Exam (no tools) — closed vs open weights. Interior points are approximate (≈); endpoints verified July 9, 2026.
How far behind is open, benchmark by benchmark?
Best generally available closed vs best open score. Shorter connector = smaller gap.
The price of the frontier
SWE-bench Verified score vs published output price per 1M tokens (log scale). Models without a published price or score are excluded, not estimated.
OpenRouter: open vs closed token share
Share of routed tokens. Endpoints reported; interior points approximate (≈).
Hugging Face downloads by lab origin
Share of open-model downloads by publishing lab's region. Endpoints from Hugging Face reporting; interior points approximate (≈).
The last 30 days
Every chart above moved because of these six events.
Anthropic ships Claude Fable 5 and Mythos 5
First Mythos-class models: 95.0% SWE-bench Verified (independently confirmed), new HLE high.
US export controls take Fable 5 and Mythos 5 offline
Order followed an Amazon researchers' jailbreak report; access suspended globally for 19 days.
Z.ai releases GLM-5.2 under MIT
First open-weight model to beat GPT-5.5 on SWE-bench Pro (62.1%); new open leader on HLE and the AA index.
OpenAI previews GPT-5.6; Mythos 5 partially restored
Sol/Terra/Luna preview begins; ~100 US critical-infrastructure orgs regain Mythos 5 access.
Commerce lifts the export-control order
Anthropic redeploys Fable 5 globally on Jul 1 with a new cybersecurity classifier; it retakes #1 on HLE.
GPT-5.6 Sol, Terra, and Luna go GA
Sol at $5/$30 with Terminal-Bench 2.1 SOTA (88.8%); HLE and SWE-bench scores not yet published.
The story in four beats
The frontier answers to regulators now
Claude Fable 5 shipped June 9, was taken offline June 12 by a US export-control order tied to a jailbreak disclosure, and returned July 1 — the first time the leading model spent 19 days dark by government order. Frontier capability is now regulated infrastructure.
GPT-5.6 landed today
Sol went GA July 9 at $5/$30 per 1M tokens with a new max reasoning effort and a subagent 'ultra mode', setting state of the art on Terminal-Bench 2.1 (88.8%). Terra matches GPT-5.5 at roughly half the price — the frontier price war is now explicit.
Open weights won the volume war
69.1% of OpenRouter tokens now route to open models; Chinese open weights hold six of the top ten slots. Anthropic runs ~12% of tokens yet captures the most revenue — volume and value have split.
The gap is release-driven, not secular
The HLE gap compressed to ~9 points in June (GLM-5.2 vs Gemini 3.1 Pro), then snapped to 17.3 when Fable 5 returned. Expect it to compress again on the next open flagship — the trend line moves in release-sized steps.
Go deeper
Model roster
GPT-5.6 Sol, Claude Fable 5, GLM-5.2, and the rest — context, pricing, modalities, and license nuance.
ExploreBenchmarks
Humanity's Last Exam, agentic coding, and graduate science — plus which benchmarks are already saturated.
ExploreEcosystem
Where tokens get routed on OpenRouter and which labs the world is downloading on Hugging Face.
Explore