Frontier model roster
Source-backed frontier and open-weight models — context, output limits, modalities, pricing, and license nuance. Verified July 9, 2026.
10 closed
9 open
| API / Focus | Modalities / Tools | Notes | ||||
|---|---|---|---|---|---|---|
Qwen3-235B-A22BAlibaba · Apr 2025 | Qwen3-235B-A22BOpen-source multilingual reasoning and coding235B total / 22B active MoE | 128K tokensOutput: not listed | Text Tools Hybrid thinking | Not listed | Open source GA Apache-2.0 | A strong Apache-2.0 model family anchor with hybrid thinking modes, 119-language coverage, and broad local deployment support.Qwen3 release blog |
Claude Fable 5Anthropic · Jun 2026 · redeployed Jul 1 | claude-fable-5Demanding reasoning and long-horizon agentic workFirst Mythos-class model; most capable GA Claude | 1M tokens Output: 128K tokens | Text Image Tools Adaptive | $10 in$50 outper 1M tokens | Proprietary GA | Current frontier leader: 53.3% on HLE (no tools) and 95.0% SWE-bench Verified (independently confirmed). Released June 9; offline June 12–30 under a US export-control order; redeployed globally July 1 with a new cybersecurity classifier.Anthropic model docs |
Claude Opus 4.8Anthropic · 2026 | claude-opus-4-8Complex reasoning, agentic coding, high-autonomy workMost capable Opus-tier model | 1M tokens (200K on Microsoft Foundry) Output: 128K tokens | Text Image Tools Adaptive | $5 in$25 outper 1M tokens | Proprietary GA | A lower-priced alternative to Fable for complex Claude workloads, with the same first-party 1M context on Claude API.Anthropic model docs |
Claude Sonnet 4.6Anthropic · Feb 2026 | claude-sonnet-4-6Balanced intelligence, latency, and priceFast frontier Claude model | 1M tokens Output: 64K tokens | Text Image Tools Extended and adaptive | $3 in$15 outper 1M tokens | Proprietary GA | The strongest Claude price-performance row for teams that need near-frontier intelligence without Opus or Fable pricing.Anthropic model docs |
DeepSeek-V4-ProDeepSeek · Apr 2026 | deepseek-v4-proOpen-weight reasoning, STEM, coding, agents1.6T total / 49B active MoE | 1M tokens Output: 384K tokens | Text Tools Thinking and non-thinking | $0.435 in$0.87 outper 1M tokensCache-miss input price shown | Open weights Preview Open weights | DeepSeek's open-weight V4 preview flagship, with 1M context, very large output budget, and strong agentic coding claims.DeepSeek API docs |
DeepSeek-V4-FlashDeepSeek · Apr 2026 | deepseek-v4-flashLow-cost open-weight reasoning284B total / 13B active MoE | 1M tokens Output: 384K tokens | Text Tools Thinking and non-thinking | $0.14 in$0.28 outper 1M tokensCache-miss input price shown | Open weights Preview Open weights | A unusually inexpensive long-context model that keeps the V4 1M context and dual thinking modes.DeepSeek API docs |
Gemini 3.1 Pro PreviewGoogle · Feb 2026 | gemini-3.1-pro-previewAgentic workflows, coding, grounded multi-step executionAdvanced Gemini 3 Pro preview | 1,048,576 tokens Output: 65.536K tokens | Text Image Video Audio PDF Tools Thinking | See Google AI pricing | Proprietary Preview | A broad multimodal Gemini preview with code execution, search grounding, URL context, structured output, and function calling.Google model docs |
Gemini 3.5 FlashGoogle · May 2026 | gemini-3.5-flashLow-latency frontier work at scaleStable high-throughput Gemini 3 model | 1,048,576 tokens Output: 65.536K tokens | Text Image Video Audio PDF Tools Thinking | See Google AI pricing | Proprietary GA | Google's stable Gemini 3 Flash row emphasizes sustained frontier performance for agentic loops, coding cycles, and high-volume workflows.Google model docs |
Llama 4 ScoutMeta · Apr 2025 | Llama-4-ScoutLong-context open-weight multimodal work17B active / 109B total MoE | 10M tokens Output: not listed | Text Image No first-party tools General | Not listed | Open weights GA Llama 4 Community License | The standout open-weight long-context option, with native multimodality, single-H100 efficiency, and a 10M context window.Meta Llama 4 docs |
Llama 4 MaverickMeta · Apr 2025 | Llama-4-MaverickOpen-weight multimodal assistant and chat use17B active / 400B total MoE | Deployment-dependentOutput: not listed | Text Image No first-party tools General | Not listed | Open weights GA Llama 4 Community License | Meta's higher-quality open-weight Llama 4 model, optimized for image and text understanding with efficient MoE inference.Meta Llama 4 blog |
Mistral Medium 3.5Mistral · Apr 2026 | mistral-medium-3-5Agentic and coding use casesFrontier-class multimodal model | 256K tokensOutput: not listed | Text Image Tools General | $1.5 in$7.5 outper 1M tokens | Open weights GA Modified MIT | Mistral's current frontier-class open-weight model, positioned for agentic and coding workloads with a 256K context window.Mistral model card |
Mistral Small 4Mistral · Mar 2026 | mistral-small-2603Efficient instruct, reasoning, and coding119B total / 6.5B active MoE | 256K tokensOutput: not listed | Text Image Tools Hybrid | $0.15 in$0.6 outper 1M tokens | Open weights GA Open weights | A strong low-cost open-weight Mistral row for high-volume use where price, latency, and 256K context matter.Mistral model card |
Mistral Large 3Mistral · Dec 2025 | mistral-large-2512Open-weight general-purpose multimodal work41B active / 675B total MoE | 256K tokensOutput: not listed | Text Image Tools General | $0.5 in$1.5 outper 1M tokens | Open weights GA Open weights | Mistral's large open-weight MoE row, useful when a larger general-purpose model is needed but 1M context is not.Mistral model card |
GPT-5.6 SolOpenAI · Jul 2026 | gpt-5.6-solDeep reasoning, agentic coding, terminal workflowsFlagship of the GPT-5.6 family (Sol / Terra / Luna) | 1.05M tokens (reported) Output: not listed | Text Image Tools Configurable incl. new max effort; ultra mode (subagents) | $5 in$30 outper 1M tokens | Proprietary GA | GA July 9, 2026. Sets state of the art on Terminal-Bench 2.1 (88.8%; 91.9% in ultra mode). Terra (~GPT-5.5 level at $2.50/$15) and Luna ($1/$6) sit below it. HLE and SWE-bench scores not yet published.OpenAI announcement |
GPT-5.5OpenAI · 2026 | gpt-5.5Complex reasoning, coding, professional workFlagship reasoning and coding model | 1M tokens Output: 128K tokens | Text Image Tools Configurable | $5 in$30 outper 1M tokens | Proprietary GA | OpenAI's previous flagship, superseded by the GPT-5.6 family on July 9, 2026 — GPT-5.6 Terra reportedly matches it at roughly half the price. Still widely deployed with 1M context and first-party tools.OpenAI model docs |
GPT-5.4 miniOpenAI · 2026 | gpt-5.4-miniCost-sensitive coding and agent workflowsLower-latency GPT-5.4 variant | 400K tokensOutput: 128K tokens | Text Image Tools Configurable | $0.75 in$4.5 outper 1M tokens | Proprietary GA | A practical lower-cost OpenAI option that keeps strong tool support, configurable reasoning, and a large output budget.OpenAI model docs |
Grok 4.3xAI · May 2026 | grok-4.3Agentic tool calling and fast general useGeneral Grok model | 1M tokens Output: not listed | Text Image Tools Configurable | $1.25 in$2.5 outper 1M tokens | Proprietary GA | xAI positions Grok 4.3 as its default general-purpose model, with 1M context, low listed pricing, and agentic tool calling.xAI model docs |
Grok Build 0.1xAI · May 2026 | grok-build-0.1Agentic coding workflowsCoding model beta | 256K tokensOutput: not listed | Text Tools Coding-specialized | $1 in$2 outper 1M tokens | Proprietary Preview | A dedicated fast coding model for agentic development loops, with a 256K context window and public beta API availability.xAI model docs |
GLM-5.2Z.ai (Zhipu) · Jun 2026 | glm-5.2Open long-horizon agentic coding and reasoning744B total / 40B active MoE | 1M tokens Output: not listed | Text Tools Thinking | $1.4 in$4.4 outper 1M tokens | Open source GA MIT | The open-weight frontier leader: first open model to beat GPT-5.5 on SWE-bench Pro (62.1%), best open scores on HLE and the AA Intelligence Index. MIT-licensed weights on Hugging Face at roughly 1/6th GPT-5.5's API cost.Z.ai release via VentureBeat |
Showing 19 of 19 models
Key takeaways
List prices span two orders of magnitude
- Output tokens range from $50/1M (Claude Fable 5) to $0.28/1M (DeepSeek-V4-Flash) — a ~180× spread inside one table. GLM-5.2 undercuts GPT-5.5 by ~6× while beating it on SWE-bench Pro.
- This spread is why OpenRouter volume flipped open (69%) while revenue stayed closed: workloads migrate to the cheapest model that clears their quality bar, not to the best model available.
The price war reached the frontier
- GPT-5.6 shipped today as three durable tiers: Sol ($5/$30), Terra (~GPT-5.5 quality at $2.50/$15), and Luna ($1/$6). Matching last quarter's flagship at half price is now the expected move.
- 1M-token context is table stakes across the frontier — Sol lists ~1.05M, and every current flagship row in this table is at 1M+. Max output budgets are where rows still differ (DeepSeek V4 lists 384K).
"Open" is three different licenses
- This table separates open source (GLM-5.2 and Qwen3 under MIT/Apache-2.0), open weights (DeepSeek, Mistral), and community-licensed (Llama 4) because they carry different commercial and redistribution rights.
- Before building on an "open" model, check the row's license note: fine-tuning, redistribution, and derivative releases are exactly where the licenses differ.
Availability is now a policy variable
- Claude Fable 5 spent June 12–30 offline under a US export-control order — the first frontier model pulled by regulators. Claude Mythos 5 remains restricted to approved organizations.
- If a model is load-bearing in your stack, plan for the possibility that its availability changes with less than a day's notice — by provider decision or by government order.