Benchmark leaderboards
How closed and open models stack up — one reporting methodology per ladder, scores from other harnesses noted rather than mixed in. Bars are colored by access so the open–closed gap reads at a glance.
Agentic Coding
SWE-bench Verified · % resolved
Share of real GitHub issues an autonomous agent resolves with a passing patch. Scores depend on the agent harness; top entries independently confirmed by vals.ai.
Graduate Science
GPQA Diamond · % correct
PhD-level multiple-choice questions in biology, physics, and chemistry, written to be Google-proof for non-experts.
Frontier Reasoning
Humanity's Last Exam (no tools) · % correct
2,500 expert-written, subject-diverse questions — the benchmark with the most headroom left at the frontier. No-tools configuration; tool-augmented runs score far higher.
Benchmarks are consumables
AIME 2026 is done as a signal — GPT-5 hit 100% and the frontier cluster sits above 95%. GPQA is inside its noise band. The discriminating benchmarks in July 2026 are HLE (no model above 54%) and Terminal-Bench 2.1, where GPT-5.6 Sol just set the state of the art at 88.8% (91.9% in ultra mode). When a benchmark saturates, retire it from your dashboards — comparisons inside the noise band are astrology.
Retired this cycle
Benchmark sources checked 2026-07-09.
Humanity's Last Exam
A Scale Labs and Center for AI Safety benchmark of 2,500 difficult, subject-diverse, multimodal questions — the frontier benchmark with the most headroom left. GPT-5.6 Sol (GA July 9) has not published an HLE score yet.
Best score (no tools)
53.3% · Claude Fable 5
Best open-weight score
≈36% · GLM-5.2
Questions
2,500
Subjects
Mathematics, Humanities, Natural Sciences, Multimodal reasoning
| Notes | |||||
|---|---|---|---|---|---|
#1 | Claude Fable 5Jun 2026 | no tools | Closed | 53.3% | Posted Jun 9; offline Jun 12–30 under US export controls; retook #1 on Jul 1 redeploy |
#2 | Claude Opus 4.82026 | no tools | Closed | 45.7% | |
#3 | Gemini 3.1 Pro PreviewFeb 2026 | no tools | Closed | 44.7% | Held #1 while Fable 5 was offline |
#4 | GPT-5.4 ProMar 2026 | pro, no tools | Closed | 44.3% | |
#5 | GPT-5.52026 | xhigh, no tools | Closed | 44.0% | |
#6 | Gemini 3 ProNov 2025 | no tools | Closed | 37.5% | |
#7 | GLM-5.2Jun 2026 | thinking, no tools | Open | ≈36% | Best open weight (MIT license); AA reports the top open cluster at 34–36% |
#8 | Kimi K2.72026 | no tools | Open | ≈35% | Top open cluster per Artificial Analysis |
#9 | DeepSeek V4 ProApr 2026 | thinking, no tools | Open | ≈34% | Top open cluster per Artificial Analysis |
#10 | GPT-5 ProOct 2025 | pro, no tools | Closed | 31.6% | |
#11 | GPT-5Aug 2025 | no tools | Closed | 25.3% | |
#12 | Kimi K2 ThinkingNov 2025 | thinking, no tools | Open | 23.9% | Open frontier at the time |
#13 | Gemini 2.5 ProJun 2025 | no tools | Closed | 21.6% | |
#14 | OpenAI o3Apr 2025 | high, no tools | Closed | 20.3% | |
#15 | DeepSeek R1-0528May 2025 | no tools | Open | 17.7% | |
#16 | DeepSeek R1Jan 2025 | no tools | Open | 8.5% | Where the open frontier stood 18 months ago |