Benchmark leaderboards

How closed and open models stack up — one reporting methodology per ladder, scores from other harnesses noted rather than mixed in. Bars are colored by access so the open–closed gap reads at a glance.

Closed / proprietary
Open weights & source
Agentic Coding: Δ 13.5 pts
Terminal Agents: Δ 16.4 pts
Frontier Reasoning: Δ 12.2 pts

Agentic Coding

SWE-bench Pro · % resolved

The harder SWE-bench successor, drawn from actively maintained repositories. Replaces SWE-bench Verified in these ladders this cycle, after Verified saturated above 95%.

A real ladder again: 30 points separate the top and bottom rows where Verified had compressed the same models into nine. Fable 5.1 leads at 81.2%, but the row that moved is Qwen3.8-Max at 67.7% — the best open score, and unlike August it is a model whose weights you can actually download. GPT-6 Astra has no cell here; OpenAI quoted DeepSWE at launch instead, so there is no like-for-like comparison with the model that leads the terminal ladder.

Terminal Agents

Terminal-Bench 4.0 · % solved

End-to-end tasks an agent completes by driving a real terminal. Version 4.0 runs five independent trials per task with an eight-hour timeout and publishes run-to-run variance. Replaces the 2.1 ladder this cycle.

The harness reset rewrote the board. GPT-6 Astra leads at 58.2% with Fable 5.1 0.3 points behind — the closest top two on the site — while Gemini 3.8 Flash, which scored 87.6% on 2.1, lands at 19.1% and Grok 4.5 falls from 83.3% to 12.4%. GLM-5.3 is the top open entry at 41.8%, ahead of Meta's closed Muse Spark 1.3. Read the drops as a statement about 2.1, not about the models: every figure above is from the same five-trial harness.

Frontier Reasoning

Humanity's Last Exam (no tools) · % correct

2,500 expert-written, subject-diverse questions — the benchmark with the most headroom left at the frontier. No-tools configuration; tool-augmented runs score far higher.

Fable 5.1 took the record to 59.1% on September 1 and GPT-6 Astra landed below three Claudes at 54.7% two days later — the flagship that leads the terminal ladder is fourth here. Kimi K3 still holds the open line at 46.9%, two months unbeaten, with the gap at 12.2 points. Note the re-run: this board's figures for unchanged models moved several points since August (Opus 5 56.3 → 54.9, Fable 5 53.3 → 55.5), so compare within this snapshot, not against the last one. Still the only headline benchmark where nobody clears 60% without tools.

Two ladders retired in one cycle

Both coding ladders on this site were replaced this month, and the reasons are different. SWE-bench Verified saturated the ordinary way: three models above 95%, the top two under a point apart, and Anthropic itself now reporting SWE-bench Pro for launches. Terminal-Bench 2.1 failed the other way — it stopped measuring what it claimed to. Gemini 3.8 Flash scored 87.6% on 2.1 and 19.1% on 4.0; Grok 4.5 went 83.3% to 12.4%. A harness that ranks a workhorse model above frontier flagships is not a hard benchmark with noise, it is a benchmark whose tasks the ecosystem has learned. Terminal-Bench 4.0 answers with five independent trials per task, an eight-hour timeout, and published variance — which is why the top scores now sit near 58% instead of near 90%. Expect the same fate: the useful question about any ladder is not who leads it but how long until the leader's score stops moving. HLE remains the exception, with nobody past 60% without tools after twenty months.

Retired

AIME (competition math)
saturated — GPT-5 at 100%
GPQA Diamond
retired Jul — noise band
SWE-bench Verified
retired Sep — three models above 95%
Terminal-Bench 2.1
retired Sep — 87.6% → 19.1% on 4.0
Promoted in their place: SWE-bench Pro, which spreads the same field 30 points, and Terminal-Bench 4.0, which runs five independent trials per task and publishes variance. Both have open-weight entries.

Benchmark sources checked 2026-09-15.

Humanity's Last Exam

Flagship reasoning benchmark

A Scale Labs and Center for AI Safety benchmark of 2,500 difficult, subject-diverse, multimodal questions — the frontier benchmark with the most headroom left. September added two flagships in three days: Claude Fable 5.1 (Sep 1) took the record to 59.1%, and GPT-6 Astra (Sep 3) landed fourth at 54.7%, behind every current Claude. The open line has not moved on a release since Kimi K3 in July.

Best score (no tools)

59.1% · Claude Fable 5.1

Best open-weight score

46.9% · Kimi K3

Questions

2,500

Subjects

Mathematics, Humanities, Natural Sciences, Multimodal reasoning

Even the leader answers under 60% of these questions — HLE is the benchmark with the most headroom left at the frontier, twenty months in. Configuration matters more than ever: this table is no-tools only, and tool-augmented runs score far higher (Fable 5.1 is reported at 65.0% with tools against 59.1% without, and GPT-6 Astra at 57.2% against 54.7%). Harness matters too — Scale's own board orders these models differently, putting Astra above Fable 5.1. The board also re-ran its current entries this month, moving scores for models that did not change; rows carrying older cells are marked in the notes column.
Notes
#1
Claude Fable 5.1Sep 2026
no tools
Closed
59.1%
Released Sep 1; reported at 65.0% with tools — a different configuration, not charted here
#2
Claude Fable 5Jun 2026
no tools
Closed
55.5%
Offline Jun 12–30 under US export controls; led from Jul 1 to Jul 24. Re-run board (was 53.3%)
#3
Claude Opus 5Jul 2026
no tools
Closed
54.9%
Held #1 from Jul 24 to Sep 1; 64.7% with tools. Re-run board (was 56.3%)
#4
GPT-6 AstraSep 2026
no tools
Closed
54.7%
Leads Terminal-Bench 4.0 but places fourth here — 57.2% with tools, still behind every current Claude
#5
GPT-5.6 SolJul 2026
no tools
Closed
47.2%
Pre-re-run cell, published Jul 27; the board has not republished it
#6
Kimi K3Jul 2026
no tools
Open
46.9%
Best open weight for two months running — 2.8T MoE; gap to the leader is 12.2 pts. Re-run board (was 43.5%)
#7
Claude Opus 4.82026
no tools
Closed
45.7%
Pre-re-run cell; superseded twice and dropped from the model roster this cycle
#8
Gemini 3.1 Pro PreviewFeb 2026
no tools
Closed
44.7%
Held #1 while Fable 5 was offline in June; still Google's most recent frontier cell seven months on
#9
GPT-5.4 ProMar 2026
pro, no tools
Closed
44.3%
#10
GPT-5.52026
xhigh, no tools
Closed
44.0%
#11
Qwen3.8-MaxAug 2026
no tools
Open
43.0%
Weights published Aug 12 under a bespoke licence — the second open model with an exact cell here
#12
GLM-5.3Aug 2026
thinking, no tools
Open
42.3%
Weights Aug 28; 6.3 pts above GLM-5.2's cluster estimate and the top open model on Terminal-Bench 4.0
#13
Gemini 3 ProNov 2025
no tools
Closed
37.5%
#14
GLM-5.2Jun 2026
thinking, no tools
Open
≈36%
Last MIT-licensed Z.ai flagship; AA put the top open cluster at 34–36% before Kimi K3 shipped
#15
Kimi K2.72026
no tools
Open
≈35%
Superseded by Kimi K3 in July
#16
DeepSeek V4 ProApr 2026
thinking, no tools
Open
≈34%
Top open cluster per Artificial Analysis; left preview Aug 13 without a published HLE cell
#17
GPT-5 ProOct 2025
pro, no tools
Closed
31.6%
#18
GPT-5Aug 2025
no tools
Closed
25.3%
#19
Kimi K2 ThinkingNov 2025
thinking, no tools
Open
23.9%
Open frontier at the time — 23 points below where Moonshot's own K3 landed eight months later
#20
Gemini 2.5 ProJun 2025
no tools
Closed
21.6%
#21
OpenAI o3Apr 2025
high, no tools
Closed
20.3%
#22
DeepSeek R1-0528May 2025
no tools
Open
17.7%
#23
DeepSeek R1Jan 2025
no tools
Open
8.5%
Where the open frontier stood 20 months ago
Showing 23 of 23 model configurations

Historical rows show scores at release to make the trajectory visible — from DeepSeek R1's 8.5% in January 2025 to Claude Fable 5.1's 59.1% twenty months later. The open line covered the same distance in less time: Moonshot went from 23.9% (K2 Thinking, November 2025) to 46.9% (K3, July 2026), and no open release has passed it in the two months since.

Data source: HLE leaderboard via public trackers · Sources checked September 15, 2026