Benchmark leaderboards

How closed and open models stack up — one reporting methodology per ladder, scores from other harnesses noted rather than mixed in. Bars are colored by access so the open–closed gap reads at a glance.

Closed / proprietary
Open weights & source
Agentic Coding: Δ 14.4 pts
Graduate Science: Δ 4.3 pts
Frontier Reasoning: Δ 17.3 pts

Agentic Coding

SWE-bench Verified · % resolved

Share of real GitHub issues an autonomous agent resolves with a passing patch. Scores depend on the agent harness; top entries independently confirmed by vals.ai.

Anthropic owns the top slots (Mythos 5 posts 95.5% but is restricted-access). The headline: DeepSeek V4 Pro Max ties Gemini 3.1 Pro at 80.6% — closed-frontier coding from an open checkpoint at roughly 3% of the price.

Graduate Science

GPQA Diamond · % correct

PhD-level multiple-choice questions in biology, physics, and chemistry, written to be Google-proof for non-experts.

A photo finish inside the noise band — GPQA is saturating. The best open model sits 4.3 pts off the lead, and this benchmark no longer separates the frontier. Watch Terminal-Bench and HLE instead.

Frontier Reasoning

Humanity's Last Exam (no tools) · % correct

2,500 expert-written, subject-diverse questions — the benchmark with the most headroom left at the frontier. No-tools configuration; tool-augmented runs score far higher.

Fable 5 retook #1 on July 1 the day it came back online. The best open cluster (GLM-5.2, Kimi K2.7, DeepSeek V4) sits at 34–36% per Artificial Analysis — about where the closed frontier stood in late 2025.

Benchmarks are consumables

AIME 2026 is done as a signal — GPT-5 hit 100% and the frontier cluster sits above 95%. GPQA is inside its noise band. The discriminating benchmarks in July 2026 are HLE (no model above 54%) and Terminal-Bench 2.1, where GPT-5.6 Sol just set the state of the art at 88.8% (91.9% in ultra mode). When a benchmark saturates, retire it from your dashboards — comparisons inside the noise band are astrology.

Retired this cycle

AIME (competition math)
saturated — GPT-5 at 100%
GPQA Diamond
inside noise band
Successor to watch: Terminal-Bench 2.1 — GPT-5.6 Sol set SOTA today at 88.8% (91.9% ultra).

Benchmark sources checked 2026-07-09.

Humanity's Last Exam

Flagship reasoning benchmark

A Scale Labs and Center for AI Safety benchmark of 2,500 difficult, subject-diverse, multimodal questions — the frontier benchmark with the most headroom left. GPT-5.6 Sol (GA July 9) has not published an HLE score yet.

Best score (no tools)

53.3% · Claude Fable 5

Best open-weight score

≈36% · GLM-5.2

Questions

2,500

Subjects

Mathematics, Humanities, Natural Sciences, Multimodal reasoning

Even the leader answers just over half of these questions — HLE is the benchmark with the most headroom left at the frontier. Configuration matters: this table is no-tools only, and tool-augmented runs score far higher (some trackers report Fable 5 at 64.5% with tools). The open rows marked ≈ come from Artificial Analysis' cluster reporting rather than exact leaderboard cells.
Notes
#1
Claude Fable 5Jun 2026
no tools
Closed
53.3%
Posted Jun 9; offline Jun 12–30 under US export controls; retook #1 on Jul 1 redeploy
#2
Claude Opus 4.82026
no tools
Closed
45.7%
#3
Gemini 3.1 Pro PreviewFeb 2026
no tools
Closed
44.7%
Held #1 while Fable 5 was offline
#4
GPT-5.4 ProMar 2026
pro, no tools
Closed
44.3%
#5
GPT-5.52026
xhigh, no tools
Closed
44.0%
#6
Gemini 3 ProNov 2025
no tools
Closed
37.5%
#7
GLM-5.2Jun 2026
thinking, no tools
Open
≈36%
Best open weight (MIT license); AA reports the top open cluster at 34–36%
#8
Kimi K2.72026
no tools
Open
≈35%
Top open cluster per Artificial Analysis
#9
DeepSeek V4 ProApr 2026
thinking, no tools
Open
≈34%
Top open cluster per Artificial Analysis
#10
GPT-5 ProOct 2025
pro, no tools
Closed
31.6%
#11
GPT-5Aug 2025
no tools
Closed
25.3%
#12
Kimi K2 ThinkingNov 2025
thinking, no tools
Open
23.9%
Open frontier at the time
#13
Gemini 2.5 ProJun 2025
no tools
Closed
21.6%
#14
OpenAI o3Apr 2025
high, no tools
Closed
20.3%
#15
DeepSeek R1-0528May 2025
no tools
Open
17.7%
#16
DeepSeek R1Jan 2025
no tools
Open
8.5%
Where the open frontier stood 18 months ago
Showing 16 of 16 model configurations

Historical rows show scores at release to make the trajectory visible — from DeepSeek R1's 8.5% in January 2025 to Fable 5's 53.3% eighteen months later.

Data source: Scale AI Leaderboard via public trackers · Sources checked July 9, 2026