Benchmark leaderboards
How closed and open models stack up — one reporting methodology per ladder, scores from other harnesses noted rather than mixed in. Bars are colored by access so the open–closed gap reads at a glance.
Agentic Coding
SWE-bench Pro · % resolved
The harder SWE-bench successor, drawn from actively maintained repositories. Replaces SWE-bench Verified in these ladders this cycle, after Verified saturated above 95%.
Terminal Agents
Terminal-Bench 4.0 · % solved
End-to-end tasks an agent completes by driving a real terminal. Version 4.0 runs five independent trials per task with an eight-hour timeout and publishes run-to-run variance. Replaces the 2.1 ladder this cycle.
Frontier Reasoning
Humanity's Last Exam (no tools) · % correct
2,500 expert-written, subject-diverse questions — the benchmark with the most headroom left at the frontier. No-tools configuration; tool-augmented runs score far higher.
Two ladders retired in one cycle
Both coding ladders on this site were replaced this month, and the reasons are different. SWE-bench Verified saturated the ordinary way: three models above 95%, the top two under a point apart, and Anthropic itself now reporting SWE-bench Pro for launches. Terminal-Bench 2.1 failed the other way — it stopped measuring what it claimed to. Gemini 3.8 Flash scored 87.6% on 2.1 and 19.1% on 4.0; Grok 4.5 went 83.3% to 12.4%. A harness that ranks a workhorse model above frontier flagships is not a hard benchmark with noise, it is a benchmark whose tasks the ecosystem has learned. Terminal-Bench 4.0 answers with five independent trials per task, an eight-hour timeout, and published variance — which is why the top scores now sit near 58% instead of near 90%. Expect the same fate: the useful question about any ladder is not who leads it but how long until the leader's score stops moving. HLE remains the exception, with nobody past 60% without tools after twenty months.
Retired
Benchmark sources checked 2026-09-15.
Humanity's Last Exam
A Scale Labs and Center for AI Safety benchmark of 2,500 difficult, subject-diverse, multimodal questions — the frontier benchmark with the most headroom left. September added two flagships in three days: Claude Fable 5.1 (Sep 1) took the record to 59.1%, and GPT-6 Astra (Sep 3) landed fourth at 54.7%, behind every current Claude. The open line has not moved on a release since Kimi K3 in July.
Best score (no tools)
59.1% · Claude Fable 5.1
Best open-weight score
46.9% · Kimi K3
Questions
2,500
Subjects
Mathematics, Humanities, Natural Sciences, Multimodal reasoning
| Notes | |||||
|---|---|---|---|---|---|
#1 | Claude Fable 5.1Sep 2026 | no tools | Closed | 59.1% | Released Sep 1; reported at 65.0% with tools — a different configuration, not charted here |
#2 | Claude Fable 5Jun 2026 | no tools | Closed | 55.5% | Offline Jun 12–30 under US export controls; led from Jul 1 to Jul 24. Re-run board (was 53.3%) |
#3 | Claude Opus 5Jul 2026 | no tools | Closed | 54.9% | Held #1 from Jul 24 to Sep 1; 64.7% with tools. Re-run board (was 56.3%) |
#4 | GPT-6 AstraSep 2026 | no tools | Closed | 54.7% | Leads Terminal-Bench 4.0 but places fourth here — 57.2% with tools, still behind every current Claude |
#5 | GPT-5.6 SolJul 2026 | no tools | Closed | 47.2% | Pre-re-run cell, published Jul 27; the board has not republished it |
#6 | Kimi K3Jul 2026 | no tools | Open | 46.9% | Best open weight for two months running — 2.8T MoE; gap to the leader is 12.2 pts. Re-run board (was 43.5%) |
#7 | Claude Opus 4.82026 | no tools | Closed | 45.7% | Pre-re-run cell; superseded twice and dropped from the model roster this cycle |
#8 | Gemini 3.1 Pro PreviewFeb 2026 | no tools | Closed | 44.7% | Held #1 while Fable 5 was offline in June; still Google's most recent frontier cell seven months on |
#9 | GPT-5.4 ProMar 2026 | pro, no tools | Closed | 44.3% | |
#10 | GPT-5.52026 | xhigh, no tools | Closed | 44.0% | |
#11 | Qwen3.8-MaxAug 2026 | no tools | Open | 43.0% | Weights published Aug 12 under a bespoke licence — the second open model with an exact cell here |
#12 | GLM-5.3Aug 2026 | thinking, no tools | Open | 42.3% | Weights Aug 28; 6.3 pts above GLM-5.2's cluster estimate and the top open model on Terminal-Bench 4.0 |
#13 | Gemini 3 ProNov 2025 | no tools | Closed | 37.5% | |
#14 | GLM-5.2Jun 2026 | thinking, no tools | Open | ≈36% | Last MIT-licensed Z.ai flagship; AA put the top open cluster at 34–36% before Kimi K3 shipped |
#15 | Kimi K2.72026 | no tools | Open | ≈35% | Superseded by Kimi K3 in July |
#16 | DeepSeek V4 ProApr 2026 | thinking, no tools | Open | ≈34% | Top open cluster per Artificial Analysis; left preview Aug 13 without a published HLE cell |
#17 | GPT-5 ProOct 2025 | pro, no tools | Closed | 31.6% | |
#18 | GPT-5Aug 2025 | no tools | Closed | 25.3% | |
#19 | Kimi K2 ThinkingNov 2025 | thinking, no tools | Open | 23.9% | Open frontier at the time — 23 points below where Moonshot's own K3 landed eight months later |
#20 | Gemini 2.5 ProJun 2025 | no tools | Closed | 21.6% | |
#21 | OpenAI o3Apr 2025 | high, no tools | Closed | 20.3% | |
#22 | DeepSeek R1-0528May 2025 | no tools | Open | 17.7% | |
#23 | DeepSeek R1Jan 2025 | no tools | Open | 8.5% | Where the open frontier stood 20 months ago |