Frontier model roster
Source-backed frontier and open-weight models — context, output limits, modalities, pricing, and license nuance. Verified September 15, 2026.
14 closed
15 open
| API / Focus | Modalities / Tools | Notes | ||||
|---|---|---|---|---|---|---|
Qwen3.8-MaxAlibaba · Aug 3, 2026 · weights Aug 12 | qwen3.8-maxFrontier reasoning, computer use, and research workflows2.4T total / 95B active MoE | 1M tokens (API) Output: not listed | Text Image Video Tools Thinking | $2 in$6 outper 1M tokensImplicit cache reads $0.25 per 1M | Open weights GA Qwen3.8-Max License — attribution at consumer scale, paid licence for MaaS and AI-assistant products above $50M revenue | The commitment landed: Qwen3.8-2.4T-A95B went up on the Hub on August 12, the first downloadable Max-tier Qwen. Read the footnotes — the published checkpoint is text-only and drops the vision and 1M window the API serves, and the licence is bespoke, not Apache-2.0. Best open score on SWE-bench Pro at 67.7% and 43.0% HLE. The Apache-2.0 release that week was the smaller Qwen3.8-27B.llm-stats open-weights analysis |
Qwen3-235B-A22BAlibaba · Apr 2025 | Qwen3-235B-A22BOpen-source multilingual reasoning and coding235B total / 22B active MoE | 128K tokensOutput: not listed | Text Tools Hybrid thinking | Not listed | Open source GA Apache-2.0 | The Apache-2.0 anchor of the family that passed 3B cumulative Hugging Face downloads with 113k+ derivatives. Alibaba's Max-tier weights finally shipped in August — but under a bespoke licence, which leaves this line, and the new Qwen3.8-27B, as the permissive option.Qwen3 release blog |
Claude Fable 5.1Anthropic · Sep 1, 2026 | claude-fable-5-1Coding, knowledge work, long-running problem solvingMythos-class flagship; Mythos 5.1 is the same model under trusted-access safeguards | 1M tokens Output: 128K tokens | Text Image PDF Tools Always-on adaptive thinking; cannot be disabled | $10 in$50 outper 1M tokensList price unchanged from Fable 5; cache reads cut 75% to $0.25 per 1M | Proprietary GA | The current frontier on two of three ladders: 59.1% HLE (no tools) and 81.2% SWE-bench Pro, with 57.9% Terminal-Bench 4.0 — 0.3 points behind GPT-6 Astra. Terminal-Bench-Science went 24.7% → 52.6% against Fable 5. Anthropic quotes SWE-bench Pro, not Verified; the 95.0% Verified figure in circulation is third-party. Same list price as Fable 5, with Anthropic putting the practical saving at ~25% on typical workloads and ~45% on agentic ones through the cache-read cut.Anthropic model docs |
Claude Opus 5Anthropic · Jul 24, 2026 | claude-opus-5Agentic coding, long-horizon autonomous work, computer useHalf-price Claude frontier tier | 1M tokens Output: 128K tokens | Text Image Tools Adaptive (on by default); effort low → max | $5 in$25 outper 1M tokens | Proprietary GA | Led every ladder on this site from July 24 until Fable 5.1 shipped on September 1. Now 54.9% HLE, 79.2% SWE-bench Pro, 51.8% Terminal-Bench 4.0 — third on the terminal ladder, six points behind the two September flagships, at half their input price and half their output price.Anthropic announcement |
Claude Fable 5Anthropic · Jun 2026 · redeployed Jul 1 | claude-fable-5Demanding reasoning and long-horizon agentic workPrevious Mythos-class flagship | 1M tokens Output: 128K tokens | Text Image Tools Always-on adaptive; effort low → max | $10 in$50 outper 1M tokens | Proprietary GA | Superseded by Fable 5.1 at the same list price but without the 75% cache-read cut — the migration is a straight discount. Still second on HLE (55.5%) and third on SWE-bench Pro (80.3%). Offline June 12–30 under a US export-control order, the first frontier model pulled by regulators.Anthropic model docs |
Claude Sonnet 5Anthropic · Jun 30, 2026 | claude-sonnet-5Balanced intelligence, latency, and priceFast frontier Claude; price-performance row | 1M tokens Output: 128K tokens | Text Image Tools Adaptive; effort low → max | $3 in$15 outper 1M tokensIntroductory $2/$10 rate expired Aug 31, 2026 — this is the standard rate | Proprietary GA | Posts 63.2% SWE-bench Pro at a fifth of Fable 5.1's output price, and remains the default model for Free and Pro users. The July introductory rate lapsed at the end of August, so the effective price of this row rose 50% while nothing about the model changed.Anthropic model docs |
DeepSeek-V4-Pro-0813DeepSeek · Apr 2026 · GA Aug 13 | deepseek-v4-proOpen-weight reasoning, STEM, coding, agents1.6T total / 49B active MoE | 1M tokens Output: 384K tokens | Text Tools Thinking and non-thinking | $0.435 in$0.87 outper 1M tokensOff-peak rate; DeepSeek has announced peak-hour output pricing up to $3.96 per 1M | Open weights GA Open weights | Left preview on August 13 as an agent-focused release, with the largest published output budget in the table at 384K. The pricing change is the story: a flat $0.87 output rate becomes up to $3.96 at peak, so the cheapest tier of the open market has started charging for congestion the way closed providers do.DeepSeek API docs |
DeepSeek-V4-FlashDeepSeek · Apr 2026 · 0731 revision | deepseek-v4-flashLow-cost open-weight reasoning at scale284B total / 13B active MoE | 1M tokens Output: 384K tokens | Text Tools Thinking and non-thinking | $0.14 in$0.28 outper 1M tokensCache-miss input price shown | Open weights GA Open weights | The most-used model on OpenRouter for the fourth month running — ~12.1T tokens a week, more than the #3 and #4 models combined. 1M context and dual thinking modes at $0.28 per 1M output is the offer the rest of the market is pricing against.DeepSeek API docs |
Gemini 3.8 FlashGoogle · Sep 2, 2026 | gemini-3.8-flashHigh-throughput agentic work, coding, document comprehensionThird Flash release in six weeks; 3.8 Flash Cyber is the gated sibling | 1,048,576 tokens Output: 65.536K tokens | Text Image Video Audio PDF Tools Thinking | $0.75 in$3.75 outper 1M tokensIntroductory rate through Dec 31, 2026; doubles to $1.50/$7.50 on Jan 1, 2027 — the end date is printed | Proprietary GA | Beats 3.7 Flash on every benchmark Google published at the same price, and beats Claude Opus 5 on three of them — including 87.6% on Terminal-Bench 2.1. Then Terminal-Bench 4.0 scored it at 19.1%, a ~68-point drop, which says more about the 2.1 harness than about this model. Google has now shipped Flash three times (Jul 21, Aug 13, Sep 2) without refreshing the Pro line.Google model docs |
Gemini 3.1 Pro PreviewGoogle · Feb 2026 | gemini-3.1-pro-previewAgentic workflows, coding, grounded multi-step executionGoogle's Pro-tier preview | 1,048,576 tokens Output: 65.536K tokens | Text Image Video Audio PDF Tools Thinking | See Google AI pricing | Proprietary Preview | Seven months old, still labelled preview, and still Google's nominal top tier while three Flash generations shipped past it. It held #1 on HLE for three weeks in June, when Fable 5 was offline under export controls; it has not had a published frontier cell since.Google model docs |
Muse Spark 1.3Meta · Sep 2, 2026 | muse-spark-1.3Agentic coding across large repositoriesCoding-first flagship behind the Muse Code agent | 1,048,576 tokens Output: not listed | Text Image Video Tools Thinking | $1.25 in$4.25 outper 1M tokensContributor endpoint at ~$0.10/$0.20 per 1M — 10–20× cheaper in exchange for letting Meta train on your traffic | Proprietary GA Open weights undecided; Meta's outstanding commitment covers 1.2, not this release | Meta's scorecard wins every coding row it lists (75.4% DeepSWE 1.1, 88.8% Terminal-Bench 2.1 in its own agent, 98.5% long-context retrieval) and it finishes work with ~20% fewer tool calls than 1.2. Independently, Terminal-Bench 4.0 puts it at 33.3% — fourth, 25 points off the lead. The contributor tier is the more interesting move: the first frontier-adjacent model priced explicitly in training data rather than dollars.llm-stats model card |
Muse Spark 1.2Meta · Aug 5, 2026 | muse-spark-1.2Agentic coding across large repositoriesPrevious coding flagship; subject of Meta's open-weight pledge | 1,048,576 tokens Output: not listed | Text Image Video Audio PDF Tools Thinking | $1.25 in$4.25 outper 1M tokensCached input $0.15 per 1M | Proprietary GA Weights promised Aug 10 'in the coming weeks' — still unpublished 36 days later | Kept in the table as the open-weight commitment nobody has been able to download yet. Meta said on August 10 it would open these weights; as of this pass there is no repository, no parameter count, and no licence. Its 82.9% Terminal-Bench 2.1 claim was run inside Meta's own agent and was never independently reproduced — and the 2.1 harness has since been retired here.VentureBeat launch coverage |
Muse Glimmer 30BMeta · Aug 10, 2026 | Muse-Glimmer-30BLocal, always-on agent workflows30B dense; 17 GB at 4-bit, runs on one consumer GPU | 128K tokensOutput: not listed | Text Image Tools General | Self-host; 24 GB cards at 4-bit, 32 GB for the dynamic quant | Open source GA Apache-2.0 | Still Meta's only published open weights and still the most permissive licence any US frontier lab has used for a current model. Leads its size class on MCP-Atlas (75.5) and DeepSearch QA (74.6) and posts 51.2% SWE-bench Pro. Five weeks after the Muse Spark pledge, this 30B model remains the whole of Meta's open release.Meta model card on Hugging Face |
Llama 4 ScoutMeta · Apr 2025 | Llama-4-ScoutLong-context open-weight multimodal work17B active / 109B total MoE | 10M tokens Output: not listed | Text Image No first-party tools General | Not listed | Open weights GA Llama 4 Community License | Still the longest published context window in the table at 10M — and, at 17 months old, the clearest illustration of how far the open field has moved past it. Meta's open line restarted on different terms: Muse Glimmer ships Apache-2.0, not the Community Licence this row carries.Meta Llama 4 docs |
MiniMax M3MiniMax · Jun 1, 2026 | minimax-m3Coding and agentic work with native multimodalityOpen-weight multimodal agentic model | 1,048,576 tokens Output: not listed | Text Image Video Tools Thinking | $0.3 in$1.2 outper 1M tokensDoubles above 512K input tokens; promotional rates run lower | Open weights GA Open weights, self-hostable | A top-ten OpenRouter model pairing a 1M window with native multimodality at a tenth of frontier pricing. It has no cell on any of the three current ladders, which is the normal state for models in this price band — the harnesses that separate them are vendor-specific.MiniMax M3 on OpenRouter |
Mistral Medium 3.5Mistral · Apr 2026 | mistral-medium-3-5Agentic and coding use casesFrontier-class multimodal model | 256K tokensOutput: not listed | Text Image Tools General | $1.5 in$7.5 outper 1M tokens | Open weights GA Modified MIT | Europe's strongest open-weight row and one of only two non-US, non-Chinese entries in the table — five months without a refresh while China shipped four open flagships and Meta shipped two.Mistral model card |
Mistral Small 4Mistral · Mar 2026 | mistral-small-2603Efficient instruct, reasoning, and coding119B total / 6.5B active MoE | 256K tokensOutput: not listed | Text Image Tools Hybrid | $0.15 in$0.6 outper 1M tokens | Open weights GA Open weights | A low-cost open-weight Mistral row for high-volume work where price, latency, and 256K context matter more than a 1M window — now priced above GLM-5.3-Flash, which is MIT, larger, and carries four times the context.Mistral model card |
Kimi K3Moonshot · Jul 16, 2026 | kimi-k3Long-horizon coding, knowledge work, agentic reasoning2.8T total-parameter MoE; largest open-weight model shipped | 1,048,576 tokens Output: not listed | Text Image Video Tools Always-on thinking (low / high / max, cannot be disabled) | $3 in$15 outper 1M tokensCached input $0.30 per 1M; flat across the full 1M window | Open weights GA Kimi K3 License (MIT-inspired, commercial gates) | Still the best open model on frontier reasoning at 46.9% HLE — 12.2 points off Fable 5.1 and four points clear of the next open row. Two months old and nothing open has passed it on that ladder. Check the licence before shipping: a $20M/12-month model-as-a-service gate and an attribution requirement above 100M MAU.Moonshot AI model card |
Nemotron 3 UltraNVIDIA · Jun 4, 2026 | nemotron-3-ultra-550b-a55bLong-running agentic workflows and open coding550B total / 55B active hybrid Mamba-Transformer MoE | 1M tokens Output: not listed | Text Tools Thinking | Self-host or third-party serving; free tier listed on OpenRouter | Open weights GA NVIDIA Open Model, Weights & Data License | The strongest US-origin open model on the router, and newly interesting for a non-technical reason: on September 3 NVIDIA agreed to buy Hugging Face for ~$12.9B, so the largest publisher of open weights and data now also owns the place almost everyone downloads them from.NVIDIA Nemotron research page |
GPT-6 AstraOpenAI · Sep 3, 2026 · GA Sep 8 | gpt-6-astraComputer use, browsing, software engineering, cybersecurity, scienceOpenAI flagship; Astra Pro on Pro, Business, and Enterprise plans | 1.05M tokens Output: 128K tokens | Text Image Tools Configurable; staged rollout through the Daybreak trusted-access program | $10 in$50 outper 1M tokensCached input $1, cache writes $12.50; above 272K input the whole request reprices to $20/$2/$25/$75. Batch and Flex at 50%; Fast mode doubles to $20/$100 and is unavailable with EU data residency | Proprietary GA | Leads Terminal-Bench 4.0 at 58.2% and OSWorld V2-Offline at 72.6% (65.7% for Sol), cutting average task time from ~75 to ~40 minutes, and holds 100% MRCR v2 retrieval to 512K. Loses the academic row: 54.7% HLE no-tools and 57.2% with tools, behind all three current Claudes. Its 99.9% ARC-AGI-3 figure comes from a specialized setup — a neutral harness puts it at 62.7%. First OpenAI model released cyber-capability-first, and the most expensive general model it has sold.OpenAI model docs |
GPT-5.6 SolOpenAI · Jul 9, 2026 | gpt-5.6-solDeep reasoning, agentic coding, terminal workflowsPrevious flagship of the GPT-5.6 family (Sol / Terra / Luna) | 1.05M tokens (reported) Output: not listed | Text Image Tools Configurable incl. max effort; ultra mode (subagents) | $5 in$30 outper 1M tokens | Proprietary GA | Superseded by Astra on September 3 and now the awkward row in OpenAI's lineup: it costs less per input token than Astra but more per output token than Claude Opus 5, and its Terminal-Bench cells are all on the retired 2.1 harness (88.8%, 91.9% ultra). Still the model behind Daybreak Blue, OpenAI's gated defender program.OpenAI model docs |
GPT-5.6 TerraOpenAI · Jul 9, 2026 | gpt-5.6-terraGeneral reasoning and coding at half flagship priceMid tier of the GPT-5.6 family | 1.05M tokens (reported) Output: not listed | Text Image Tools Configurable | $2.5 in$15 outper 1M tokens | Proprietary GA | Positioned as roughly GPT-5.5-level quality at half the price. With Astra listing at $10/$50, Terra is now a quarter of OpenAI's own flagship input price — the spread inside a single vendor's lineup is wider than the spread between vendors was a year ago.OpenAI model docs |
GPT-5.6 LunaOpenAI · Jul 9, 2026 | gpt-5.6-lunaHigh-volume, cost-sensitive workloadsCost tier of the GPT-5.6 family | 1.05M tokens Output: 128K tokens | Text Image Tools Extended reasoning, configurable | $0.2 in$1.2 outper 1M tokensCached input $0.02 per 1M | Proprietary GA | The only US closed model near the top of OpenRouter by volume — #3 at ~9.5T tokens a week, up 129% week over week. A 1.05M window, vision, and tool calling at $1.20 per 1M output is what it takes for a closed model to compete with open weights on bulk work. It scores 17.3% on Terminal-Bench 4.0, which is the tradeoff.OpenAI model docs |
Grok 4.6xAI · Aug 12, 2026 | grok-4.6Long-running agents, agentic coding, interactive and visual workCurrent xAI flagship | 500K tokensOutput: not listed | Text Image Tools Configurable | $2 in$6 outper 1M tokensCached input $0.50 per 1M; above 200K prompt tokens every token in the request reprices to $4 / $1 / $12 | Proprietary GA | Reached the top of the Artificial Analysis index on launch at Grok 4.5's price — but Artificial Analysis has re-based the index twice since (v4.2 on Sep 4, v4.3 on Sep 7), so that reading is no longer comparable to anything. Five weeks on it still has no HLE, SWE-bench Pro, or Terminal-Bench 4.0 cell, so it appears on no ladder here. Its published numbers (DeepSWE 1.1 65.9, APEX-Agents 57.5, GDPVal-AA v2 1,753) share no harness with the rest of the table.xAI announcement |
Grok 4.5xAI · Jul 8, 2026 | grok-4.5Agentic coding and office knowledge workPrevious xAI flagship | 500K tokensOutput: not listed | Text Image Tools Configurable | $2 in$6 outper 1M tokensCached input $0.50 per 1M; >200K-token requests can price higher | Proprietary GA | Kept here because it is the only xAI row with independent runs on the harnesses this site charts: 64.7% SWE-bench Pro and 12.4% Terminal-Bench 4.0. That second number is the useful one — on the current terminal harness the whole mid-field collapses, and a model that scored 83.3% on Terminal-Bench 2.1 lands in single-digit-to-low-teens territory.xAI model docs |
MiMo-V2.5Xiaomi · Apr 2026 | mimo-v2.5High-volume multimodal inference at the price floor310B total / 15B active MoE | 1.05M tokens Output: not listed | Text Image Audio Tools General | $0.112 in$0.224 outper 1M tokensRepresentative OpenRouter serving price | Open source GA Fully open-sourced weights | Held #1 on OpenRouter by routed tokens in late July; now #4 at ~7.2T tokens a week as the two newest cheap models passed it. A 15B-active model with native vision and audio at roughly $0.22 per 1M output — the price floor that everything else on the router is measured against.Xiaomi MiMo model card |
GLM-5.3Z.ai (Zhipu) · Aug 14, 2026 · weights Aug 28 | glm-5.3Open long-horizon agentic coding, CLI work, and security753B total / 40B active MoE (same base as GLM-5.2) | 1M tokens Output: 128K tokens | Text Tools Thinking | $1.4 in$4.4 outper 1M tokensCached input $0.26 per 1M | Open weights GA Bespoke GLM-5.3 licence — Z.ai security review required for MaaS operators above $10B revenue | The top open-weight model on Terminal-Bench 4.0 at 41.8% — 16 points behind the closed leaders and 8 points ahead of Meta's Muse Spark 1.3. Same base as 5.2 with scaled post-training: DeepSWE v1.1 66.9%, CyberGym 77.2% → 84.5%. The licence is the change to note: where GLM-5.2 shipped MIT, the 5.3 flagship carries a review clause, and MIT moved down to the Flash tier.BenchLM model page |
GLM-5.3-FlashZ.ai (Zhipu) · Aug 2026 | glm-5.3-flashHigh-volume agentic and coding work at the price floor320B total / 18B active MoE | 1M tokens Output: 128K tokens | Text Tools Thinking | $0.15 in$0.5 outper 1M tokensA 50% launch discount ($0.07/$0.25) ran to Sep 9, 2026 | Open source GA MIT | Went from launch to #2 on OpenRouter at ~10T tokens a week — the fastest adoption curve on the router this year. It is also now the strongest genuinely MIT-licensed model available, because the flagship line moved to a bespoke licence: no revenue gate, no attribution clause, no security review.GLM-5.3-Flash on OpenRouter |
GLM-5.2Z.ai (Zhipu) · Jun 13, 2026 | glm-5.2Open long-horizon agentic coding and reasoning744B total / 40B active MoE | 1M tokens Output: not listed | Text Tools Thinking | $1.4 in$4.4 outper 1M tokens | Open source GA MIT | The last MIT-licensed Z.ai flagship, kept as the comparison point for what the licence change cost: 62.1% SWE-bench Pro, no revenue gate, no attribution clause, no review requirement. Anyone who standardized on 5.2 for its licence has to re-read the terms before moving to 5.3.Z.ai release via VentureBeat |
Showing 29 of 29 models
Key takeaways
List prices span two orders of magnitude
- Output tokens range from $50/1M (Claude Fable 5.1, GPT-6 Astra) to $0.22/1M (MiMo-V2.5) — a ~220× spread inside one table. On the harder SWE-bench Pro harness the premium buys 13.5 points over the best open row: Qwen3.8-Max resolves 67.7% at $6/1M against Fable 5.1's 81.2% at $50.
- This spread is why OpenRouter volume flipped open (~69%) while revenue stayed closed: workloads migrate to the cheapest model that clears their quality bar, not to the best model available. What changed this month is the bottom, not the top — Meta will sell Muse Spark 1.3 at roughly $0.10/$0.20 if you let it train on your traffic, and DeepSeek's V4-Pro now charges up to $3.96/1M at peak against a flat $0.87.
The price war reached the frontier
- It stopped falling this month. Fable 5.1 held Fable 5's $10/$50 and discounted the cache instead (reads down 75%, ~25% cheaper in practice), and GPT-6 Astra listed at the same $10/$50 — the most expensive general model OpenAI has sold, doubling to $20/$100 above 272K input tokens. Claude Sonnet 5's introductory rate simply lapsed on August 31, raising that row 50% with no model change.
- 1M-token context is table stakes — every current flagship row here is at 1M+ except the Grok pair, which ship 500K. Rows still differ on max output budgets (DeepSeek V4 lists 384K), on whether thinking can be switched off at all (Kimi K3 and the Fable line: no), on how long-context requests are billed (Grok 4.6 reprices the whole request above 200K, Astra above 272K), and on when the price expires — Gemini 3.8 Flash prints its own doubling date of January 1, 2027.
"Open" is three different licenses
- This table separates open source (GLM-5.2, GLM-5.3-Flash, Qwen3, MiMo-V2.5, Muse Glimmer under MIT/Apache-2.0/fully-open terms), open weights (Qwen3.8-Max, GLM-5.3, Kimi K3, DeepSeek, Mistral, Nemotron), and community-licensed (Llama 4) because they carry different commercial and redistribution rights.
- Two of August's three promises were kept, and both arrived narrower than announced. Alibaba published Qwen3.8-Max's weights on August 12 — text-only, without the API's vision or 1M window, under a licence with a $50M revenue gate. Z.ai published GLM-5.3's on August 28 under a licence requiring a security review above $10B revenue, where GLM-5.2 had been MIT. Meta's Muse Spark 1.2 weights, due "in the coming weeks" on August 10, still do not exist; that row stays proprietary until they do.
- The pattern is the finding: every current open flagship now ships under bespoke terms, and the genuinely permissive rows have moved down-tier to GLM-5.3-Flash (MIT) and Muse Glimmer 30B (Apache-2.0). Kimi K3 gates model-as-a-service above $20M revenue and requires attribution above 100M MAU; Qwen3.8-Max and GLM-5.3 each add their own thresholds. Picking an open model is now a licence-review task before it is a benchmark one.
Availability is now a policy variable
- Claude Fable 5 spent June 12–30 offline under a US export-control order — the first frontier model pulled by regulators. The gating is now built into launches rather than imposed after them: Mythos 5.1 ships only through trusted-access programs, GPT-6 Astra went to OpenAI's Daybreak cyber program days before general availability, and Google shipped a Fairwind-gated Gemini 3.8 Flash Cyber alongside the public model. Beijing is separately weighing export controls on Chinese frontier and open-weight models.
- If a model is load-bearing in your stack, plan for the possibility that its availability changes with less than a day's notice — by provider decision or by government order, on either side. Open weights you have already downloaded are the one row in this table that policy cannot retract.