AI Models
28Find the right model for your task. Filter by strength, compare pricing, and decide.
Filter by provider
Claude Fable 5
Anthropic
Anthropic's Mythos-class frontier model — the strongest publicly-available Claude yet, leading every major benchmark at launch. The verdict stays CONDITIONAL on two live constraints: credit-only pricing at $10/$50 per 1M tokens with no subscription path, and a safety classifier that over-triggers on routine coding tasks, silently routing them to an Opus 4.8 fallback. Reach for Fable 5 when the capability delta justifies the cost; for routine work, Sonnet 5 is the more predictable alternative.
- Input
- $10 / 1M tokens
- Output
- $50 / 1M tokens
- Context
- 1M context, 128K max output
Claude Haiku 4.5
Anthropic
Anthropic's value baseline at $1/$5 per 1M tokens — the fast, low-cost tier rivals benchmark against, with 200K context and 64K output. Best for high-volume latency-sensitive work: classification, extraction, summarization, routing. Reach for Opus 4.8 or Sonnet 5 when deep multi-step reasoning is needed.
- Input
- $1 / 1M tokens
- Output
- $5 / 1M tokens
- Context
- 200K tokens
Claude Opus 4.8
Anthropic
Anthropic's flagship reasoning model and the practical frontier incumbent for deployable capability — 1M-token context at $5/$25 per 1M tokens, included across all subscription plans. Our pick for review-heavy work: code review, constrained refactoring, and deliberate-paced tasks. Strong agentic coding, top-tier analytical reasoning, and stable subscription access make it the default flagship for most production teams.
- Input
- $5 / 1M tokens
- Output
- $25 / 1M tokens
- Context
- 1M tokens
Claude Sonnet 4.6
Anthropic
Superseded by Claude Sonnet 5 (shipped Jun 30, 2026) as the default mid-tier workhorse. Sonnet 4.6's $3/$15 pricing and 1M-token context are now matched by a successor with a 13.4-point Terminal-Bench 2.1 jump at the same rate card. Choose 4.6 only for pinned deployments or where Sonnet 5's ~30% higher token count and removed temperature/top_p/top_k are disqualifying.
- Input
- $3 / 1M tokens
- Output
- $15 / 1M tokens
- Context
- 1M tokens
Claude Sonnet 5
Anthropic
Anthropic's most agentic Sonnet delivers near-Opus 4.8 performance at roughly half the price, with Terminal-Bench 2.1 +13.4 points over Sonnet 4.6. The new tokenizer inflates counts ~30%, so intro pricing ($2/$10 per MTok through Aug 31) keeps migration cost-neutral, but teams should recount budgets before switching. It replaces Sonnet 4.6 as the default across all Claude plans, ideal for agentic coding, brownfield debugging, and production pipelines — not for the hardest frontier reasoning or cybersecurity tasks.
- Input
- $3 / 1M
- Output
- $15 / 1M
- Context
- 1M tokens
GLM-5.2
Z.ai
Top open-weight coding model with 1M-token context and MIT license — beats GPT-5.5 on multiple benchmarks at a fraction of cost, but still trails Claude Opus 4.8 on the hardest long-horizon tasks.
- Input
- $1.40 / 1M tokens
- Output
- $4.40 / 1M tokens
- Context
- 1M tokens
GPT-4.1
OpenAI
The default migration target for GPT-4 Turbo workloads: same request shape, lower cost, 1M-token context — convert legacy function_call payloads to the tools format and move on.
- Input
- $2 / 1M tokens
- Output
- $8 / 1M tokens
- Context
- 1M tokens
GPT-4o (Retired from ChatGPT, API-only)
OpenAI
GPT-4o (May 2024) has been superseded by the GPT-5.x family and delisted from OpenAI's flagship API pricing page. It remains available through the API but is no longer a current-generation model. Use only where migration to GPT-5.x is blocked; otherwise, GPT-5.5 or GPT-5.4 offer better performance at comparable cost.
- Input
- $2.50 / 1M tokens
- Output
- $10 / 1M tokens
- Context
- 128K tokens
GPT-5.5
OpenAI
OpenAI's general-purpose commercial frontier model and the deployable OpenAI anchor while the GPT-5.6 family stays government-gated — $5/$30 per 1M tokens with 1M+ context. Strong reasoning, broad ecosystem support, and full public availability make it a solid production default, though cheaper models now match it on specific tasks. For most teams, it remains the baseline newer models are measured against.
- Input
- $5 / 1M tokens
- Output
- $30 / 1M tokens
- Context
- 1M+ tokens (922K in / 128K out)
GPT-5.5-Cyber
OpenAI
85.6% CyberGym — top single-model score — but restricted-access for vetted defenders only.
- Input
- N/A
- Output
- N/A
- Context
- Trusted Access for Cyber; not commercial API
GPT-5.6 Luna
OpenAI
OpenAI's budget GPT-5.6 tier — TerminalBench 2.1 at 82.5% for $1/$6 per 1M tokens, undercutting Claude Haiku 4.5 and Gemini Flash on capability-per-dollar. The caveat: Luna is government-gated to ~20 trusted partners, with no independent verification or GA date. Best for teams building model-routing architectures who can escalate cheap failures to Terra or Sol.
- Input
- $1 / 1M tokens
- Output
- $6 / 1M tokens
- Context
- Fastest and most cost-efficient GPT-5.6 tier.
GPT-5.6 Sol
OpenAI
GPT-5.6 Sol is the most cost-effective frontier model, leading on coding benchmarks and competitive on hard reasoning at a fraction of the cost of peers. Verdict reinforced by GPT-Red safety-hardening disclosure — the model was adversarially trained against a dedicated self-play RL red-teaming model at frontier scale, making it 6x more robust to prompt injections.
- Input
- $5 / 1M tokens
- Output
- $30 / 1M tokens
- Context
- Same rate card as GPT-5.5. Ultra mode costs more.
GPT-5.6 Terra
OpenAI
The mid-tier GPT-5.6 workhorse — GPT-5.5-class capability at half the cost ($2.50/$15 per 1M tokens), with TerminalBench 2.1 at 84.3% tying Claude Fable 5. The caveat: all claims are vendor-reported, Terra is government-gated to ~20 partners, and GA has no confirmed date. For most production workloads Terra would be the pragmatic default once available; GPT-5.5 and Opus 4.8 are the fallbacks today.
- Input
- $2.50 / 1M tokens
- Output
- $15 / 1M tokens
- Context
- Roughly 2x cheaper than GPT-5.5 for similar perf.
GPT-Live-1
OpenAI
OpenAI's first full-duplex voice model — listens and speaks simultaneously, a leap beyond turn-based Advanced Voice Mode. Delegates complex reasoning to GPT-5.5 in the background while maintaining conversational flow. Verdict CONDITIONAL: it's the best ChatGPT voice experience available, but it's a product-integrated model with no API and uneven language support.
- Input
- ChatGPT subscription
- Output
- ChatGPT subscription
- Context
- ChatGPT Voice only
GPT-Live-1 mini
OpenAI
The free-tier variant of GPT-Live-1 — same full-duplex architecture in a smaller, faster package, serving as the default voice model for ChatGPT Free users. Shares the core innovations: continuous interaction, wake word support, and GPT-5.5 delegation. Same limitations apply: no API, uneven language quality. A substantial upgrade over the old turn-based voice mode for free users.
- Input
- Free (ChatGPT)
- Output
- Free (ChatGPT)
- Context
- ChatGPT Voice only
Gemini 2.5 Pro
Now the budget long-context pick: a 40% price cut (Jun 2026) made full-corpus passes affordable at mid-size budgets — the longest context window available with strong Google ecosystem integration.
- Input
- $1.25 / 1M tokens
- Output
- $10 / 1M tokens
- Context
- 1M tokens
Gemini 3.5 Flash
Google's cost-efficiency standout for agentic desktop automation at $1.50/$9 per 1M tokens with ~1M context. Native Computer Use scores 78.4 on OSWorld-Verified — within 0.3 points of GPT-5.5 at roughly one-third the cost. THE relevant caveat: the Computer Use tool is a public preview not yet in the production API, so verify endpoint availability before wiring in.
- Input
- $1.50 / 1M tokens
- Output
- $9 / 1M tokens
- Context
- 1M tokens
Gemini 3.5 Pro
Google's pre-GA frontier model — now on its fourth postponement and 'months behind schedule' per Bloomberg. The 2M-token context, native multimodal pipeline, and Deep Think reasoning remain real capabilities, but no confirmed GA timeline exists. Coding capability is the core shortfall.
- Input
- Not yet published
- Output
- Not yet published
- Context
- Est. $5-15/$30-60 per 1M. 2M ctx. Not GA.
Grok 4.5
SpaceXAI
Competitively priced coding/agentic model at $2/$6 per 1M tokens with 500K context — xAI's published benchmarks place it mid-pack behind Fable 5 and GPT-5.5, though it leads Opus 4.8 on DeepSWE and Terminal-Bench 2.1. The real story is cost efficiency: 4.2x more token-efficient than Opus 4.8 on SWE-Bench Pro. Best for teams needing strong coding at a fraction of frontier pricing. Not yet available in the EU.
- Input
- $2 / 1M tokens
- Output
- $6 / 1M tokens
- Context
- 500K tokens
Hy3
Tencent
Apache 2.0 open-weight MoE agentic model (295B total, 21B active) with strong tool-calling reliability and anti-hallucination behavior, but coding trails GLM-5.2 and self-hosting demands 8+ GPUs. Free OpenRouter tier through Jul 21 — afterward $0.20/$0.80 per 1M tokens. Best for self-hosted coding agents and long-context document work where permissive licensing matters more than absolute benchmark scores.
- Input
- $0.20 / 1M
- Output
- $0.80 / 1M
- Context
- 256K tokens
Inkling
Thinking Machines Lab
Inkling is the strongest reason yet to take open-weight multimodal AI seriously. It delivers competitive reasoning and coding with best-in-class open-weight safety and controllable thinking effort, all under Apache 2.0. But Thinking Machines is honest that it's not the strongest overall — factuality lags (SimpleQA 43.9%) and self-hosting demands serious hardware (2TB VRAM). Conditional: a compelling base for organizations that need to own and fine-tune a multimodal model, not for teams wanting peak off-the-shelf performance.
- Input
- $1.87 / 1M tokens
- Output
- $4.68 / 1M tokens
- Context
- Up to 1M tokens native
Kimi K3
Moonshot AI
Moonshot AI's 2.8T MoE flagship — the largest open-weight model ever released — leads Arena WebDev front-end coding and scores competitively with Opus 4.8 on long-horizon coding tasks. Flat $3/$15 per 1M pricing keeps it accessible, but it still trails Fable 5 and GPT 5.6 Sol on precise-execution benchmarks, and weights remain hosted-only until July 27. Best for teams doing long-context agentic coding who can tolerate verbosity and don't need immediate self-hosting.
- Input
- $3.00 / 1M
- Output
- $15.00 / 1M
- Context
- 1M tokens, $0.30/1M cached
Llama 4 Maverick
Meta
The leading open-source model for teams that need self-hosting, data sovereignty, or want to avoid API vendor lock-in.
- Input
- Free (self-hosted) or $0.20 / 1M tokens (hosted)
- Output
- Free (self-hosted) or $0.60 / 1M tokens (hosted)
- Context
- 10M tokens
MAI-Code-1-Flash
Microsoft
Microsoft's first in-house coding model — a small, fast, aggressively cost-optimized sparse MoE at $0.75/$4.50 per 1M tokens. Delivers a 16-point lead over Claude Haiku 4.5 on SWE-Bench Pro while using up to 60% fewer tokens. It is purpose-built for high-volume Copilot workflows — autocompletions, refactors, scaffolds — not a frontier model. Best inside the Copilot ecosystem; API access outside it is still rolling out.
- Input
- $0.75 / 1M tokens
- Output
- $4.50 / 1M tokens
- Context
- 256K context
Muse Spark 1.1
Meta
Meta's first paid AI model and first salvo in the commercial coding market — $1.25/$4.25 per 1M tokens with 1M context, native subagent orchestration, MCP/tool support, and computer-use capability at roughly 4x under Anthropic/OpenAI pricing. Leads on tool-use benchmarks but trails Opus 4.8 on pure coding by 7 points on SWE-Bench Pro. Closed-weight, US-only public preview. Best for teams prioritizing tool-use orchestration at low cost.
- Input
- $1.25 / 1M tokens
- Output
- $4.25 / 1M tokens
- Context
- 1M tokens
Mythos 5 (Claude Mythos 5)
Anthropic
The strongest cybersecurity model in the world (83.8% CyberGym) — built on Fable 5's weights with safety classifiers lifted. Access was suspended after the Jun 12 export directive; partially restored Jun 26 to ~100 designated US critical-infrastructure organizations (Project Glasswing). Verdict stays CONDITIONAL: capability is elite, but access is gated by federal designation. Not a public API — usable only if your organization is on the Commerce Department list.
- Input
- $10 / 1M tokens
- Output
- $50 / 1M tokens
- Context
- 1M tokens (restricted access)
Ornith-1.0
DeepReinforce
A strong MIT-licensed open-source agentic coding model family that achieves frontier-competitive benchmarks — beating Claude Opus 4.7 on both Terminal-Bench 2.1 and SWE-Bench Verified — but requires self-hosting with significant GPU resources and comes from a new lab with limited track record.
- Input
- Free (MIT license)
- Output
- Free (MIT license)
- Context
- 256K tokens
Sakana Fugu
Sakana AI
Multi-agent orchestrator that matches frontier models on key benchmarks by routing across a pool of expert agents — but opaque routing and platform dependency limit transparency.
- Input
- $5 / 1M tokens
- Output
- $30 / 1M tokens
- Context
- 272K tokens (>272K: $10/$45 per 1M)