AI Models

28

Find the right model for your task. Filter by strength, compare pricing, and decide.

Filter by provider

28models
Claude Fable 5 logo

Claude Fable 5

Anthropic

Conditional

Anthropic's Mythos-class frontier model — the strongest publicly-available Claude yet, leading every major benchmark at launch. The verdict stays CONDITIONAL on two live constraints: credit-only pricing at $10/$50 per 1M tokens with no subscription path, and a safety classifier that over-triggers on routine coding tasks, silently routing them to an Opus 4.8 fallback. Reach for Fable 5 when the capability delta justifies the cost; for routine work, Sonnet 5 is the more predictable alternative.

Input
$10 / 1M tokens
Output
$50 / 1M tokens
Context
1M context, 128K max output
Claude Haiku 4.5 logo

Claude Haiku 4.5

Anthropic

Recommended

Anthropic's value baseline at $1/$5 per 1M tokens — the fast, low-cost tier rivals benchmark against, with 200K context and 64K output. Best for high-volume latency-sensitive work: classification, extraction, summarization, routing. Reach for Opus 4.8 or Sonnet 5 when deep multi-step reasoning is needed.

Input
$1 / 1M tokens
Output
$5 / 1M tokens
Context
200K tokens
Claude Opus 4.8 logo

Claude Opus 4.8

Anthropic

Recommended

Anthropic's flagship reasoning model and the practical frontier incumbent for deployable capability — 1M-token context at $5/$25 per 1M tokens, included across all subscription plans. Our pick for review-heavy work: code review, constrained refactoring, and deliberate-paced tasks. Strong agentic coding, top-tier analytical reasoning, and stable subscription access make it the default flagship for most production teams.

Input
$5 / 1M tokens
Output
$25 / 1M tokens
Context
1M tokens
Claude Sonnet 4.6 logo

Claude Sonnet 4.6

Anthropic

Conditional

Superseded by Claude Sonnet 5 (shipped Jun 30, 2026) as the default mid-tier workhorse. Sonnet 4.6's $3/$15 pricing and 1M-token context are now matched by a successor with a 13.4-point Terminal-Bench 2.1 jump at the same rate card. Choose 4.6 only for pinned deployments or where Sonnet 5's ~30% higher token count and removed temperature/top_p/top_k are disqualifying.

Input
$3 / 1M tokens
Output
$15 / 1M tokens
Context
1M tokens
Claude Sonnet 5 logo

Claude Sonnet 5

Anthropic

Recommended

Anthropic's most agentic Sonnet delivers near-Opus 4.8 performance at roughly half the price, with Terminal-Bench 2.1 +13.4 points over Sonnet 4.6. The new tokenizer inflates counts ~30%, so intro pricing ($2/$10 per MTok through Aug 31) keeps migration cost-neutral, but teams should recount budgets before switching. It replaces Sonnet 4.6 as the default across all Claude plans, ideal for agentic coding, brownfield debugging, and production pipelines — not for the hardest frontier reasoning or cybersecurity tasks.

Input
$3 / 1M
Output
$15 / 1M
Context
1M tokens
G

GLM-5.2

Z.ai

Conditional

Top open-weight coding model with 1M-token context and MIT license — beats GPT-5.5 on multiple benchmarks at a fraction of cost, but still trails Claude Opus 4.8 on the hardest long-horizon tasks.

Input
$1.40 / 1M tokens
Output
$4.40 / 1M tokens
Context
1M tokens
GPT-4.1 logo

GPT-4.1

OpenAI

Recommended

The default migration target for GPT-4 Turbo workloads: same request shape, lower cost, 1M-token context — convert legacy function_call payloads to the tools format and move on.

Input
$2 / 1M tokens
Output
$8 / 1M tokens
Context
1M tokens
GPT-4o (Retired from ChatGPT, API-only) logo

GPT-4o (Retired from ChatGPT, API-only)

OpenAI

Conditional

GPT-4o (May 2024) has been superseded by the GPT-5.x family and delisted from OpenAI's flagship API pricing page. It remains available through the API but is no longer a current-generation model. Use only where migration to GPT-5.x is blocked; otherwise, GPT-5.5 or GPT-5.4 offer better performance at comparable cost.

Input
$2.50 / 1M tokens
Output
$10 / 1M tokens
Context
128K tokens
GPT-5.5 logo

GPT-5.5

OpenAI

Recommended

OpenAI's general-purpose commercial frontier model and the deployable OpenAI anchor while the GPT-5.6 family stays government-gated — $5/$30 per 1M tokens with 1M+ context. Strong reasoning, broad ecosystem support, and full public availability make it a solid production default, though cheaper models now match it on specific tasks. For most teams, it remains the baseline newer models are measured against.

Input
$5 / 1M tokens
Output
$30 / 1M tokens
Context
1M+ tokens (922K in / 128K out)
GPT-5.5-Cyber logo

GPT-5.5-Cyber

OpenAI

Conditional

85.6% CyberGym — top single-model score — but restricted-access for vetted defenders only.

Input
N/A
Output
N/A
Context
Trusted Access for Cyber; not commercial API
GPT-5.6 Luna logo

GPT-5.6 Luna

OpenAI

pending

OpenAI's budget GPT-5.6 tier — TerminalBench 2.1 at 82.5% for $1/$6 per 1M tokens, undercutting Claude Haiku 4.5 and Gemini Flash on capability-per-dollar. The caveat: Luna is government-gated to ~20 trusted partners, with no independent verification or GA date. Best for teams building model-routing architectures who can escalate cheap failures to Terra or Sol.

Input
$1 / 1M tokens
Output
$6 / 1M tokens
Context
Fastest and most cost-efficient GPT-5.6 tier.
GPT-5.6 Sol logo

GPT-5.6 Sol

OpenAI

Recommended

GPT-5.6 Sol is the most cost-effective frontier model, leading on coding benchmarks and competitive on hard reasoning at a fraction of the cost of peers. Verdict reinforced by GPT-Red safety-hardening disclosure — the model was adversarially trained against a dedicated self-play RL red-teaming model at frontier scale, making it 6x more robust to prompt injections.

Input
$5 / 1M tokens
Output
$30 / 1M tokens
Context
Same rate card as GPT-5.5. Ultra mode costs more.
GPT-5.6 Terra logo

GPT-5.6 Terra

OpenAI

pending

The mid-tier GPT-5.6 workhorse — GPT-5.5-class capability at half the cost ($2.50/$15 per 1M tokens), with TerminalBench 2.1 at 84.3% tying Claude Fable 5. The caveat: all claims are vendor-reported, Terra is government-gated to ~20 partners, and GA has no confirmed date. For most production workloads Terra would be the pragmatic default once available; GPT-5.5 and Opus 4.8 are the fallbacks today.

Input
$2.50 / 1M tokens
Output
$15 / 1M tokens
Context
Roughly 2x cheaper than GPT-5.5 for similar perf.
GPT-Live-1 logo

GPT-Live-1

OpenAI

Conditional

OpenAI's first full-duplex voice model — listens and speaks simultaneously, a leap beyond turn-based Advanced Voice Mode. Delegates complex reasoning to GPT-5.5 in the background while maintaining conversational flow. Verdict CONDITIONAL: it's the best ChatGPT voice experience available, but it's a product-integrated model with no API and uneven language support.

Input
ChatGPT subscription
Output
ChatGPT subscription
Context
ChatGPT Voice only
GPT-Live-1 mini logo

GPT-Live-1 mini

OpenAI

Conditional

The free-tier variant of GPT-Live-1 — same full-duplex architecture in a smaller, faster package, serving as the default voice model for ChatGPT Free users. Shares the core innovations: continuous interaction, wake word support, and GPT-5.5 delegation. Same limitations apply: no API, uneven language quality. A substantial upgrade over the old turn-based voice mode for free users.

Input
Free (ChatGPT)
Output
Free (ChatGPT)
Context
ChatGPT Voice only
Gemini 2.5 Pro logo

Gemini 2.5 Pro

Google

Recommended

Now the budget long-context pick: a 40% price cut (Jun 2026) made full-corpus passes affordable at mid-size budgets — the longest context window available with strong Google ecosystem integration.

Input
$1.25 / 1M tokens
Output
$10 / 1M tokens
Context
1M tokens
Gemini 3.5 Flash logo

Gemini 3.5 Flash

Google

Recommended

Google's cost-efficiency standout for agentic desktop automation at $1.50/$9 per 1M tokens with ~1M context. Native Computer Use scores 78.4 on OSWorld-Verified — within 0.3 points of GPT-5.5 at roughly one-third the cost. THE relevant caveat: the Computer Use tool is a public preview not yet in the production API, so verify endpoint availability before wiring in.

Input
$1.50 / 1M tokens
Output
$9 / 1M tokens
Context
1M tokens
Gemini 3.5 Pro logo

Gemini 3.5 Pro

Google

pending

Google's pre-GA frontier model — now on its fourth postponement and 'months behind schedule' per Bloomberg. The 2M-token context, native multimodal pipeline, and Deep Think reasoning remain real capabilities, but no confirmed GA timeline exists. Coding capability is the core shortfall.

Input
Not yet published
Output
Not yet published
Context
Est. $5-15/$30-60 per 1M. 2M ctx. Not GA.
G

Grok 4.5

SpaceXAI

Conditional

Competitively priced coding/agentic model at $2/$6 per 1M tokens with 500K context — xAI's published benchmarks place it mid-pack behind Fable 5 and GPT-5.5, though it leads Opus 4.8 on DeepSWE and Terminal-Bench 2.1. The real story is cost efficiency: 4.2x more token-efficient than Opus 4.8 on SWE-Bench Pro. Best for teams needing strong coding at a fraction of frontier pricing. Not yet available in the EU.

Input
$2 / 1M tokens
Output
$6 / 1M tokens
Context
500K tokens
H

Hy3

Tencent

Conditional

Apache 2.0 open-weight MoE agentic model (295B total, 21B active) with strong tool-calling reliability and anti-hallucination behavior, but coding trails GLM-5.2 and self-hosting demands 8+ GPUs. Free OpenRouter tier through Jul 21 — afterward $0.20/$0.80 per 1M tokens. Best for self-hosted coding agents and long-context document work where permissive licensing matters more than absolute benchmark scores.

Input
$0.20 / 1M
Output
$0.80 / 1M
Context
256K tokens
I

Inkling

Thinking Machines Lab

Conditional

Inkling is the strongest reason yet to take open-weight multimodal AI seriously. It delivers competitive reasoning and coding with best-in-class open-weight safety and controllable thinking effort, all under Apache 2.0. But Thinking Machines is honest that it's not the strongest overall — factuality lags (SimpleQA 43.9%) and self-hosting demands serious hardware (2TB VRAM). Conditional: a compelling base for organizations that need to own and fine-tune a multimodal model, not for teams wanting peak off-the-shelf performance.

Input
$1.87 / 1M tokens
Output
$4.68 / 1M tokens
Context
Up to 1M tokens native
K

Kimi K3

Moonshot AI

Conditional

Moonshot AI's 2.8T MoE flagship — the largest open-weight model ever released — leads Arena WebDev front-end coding and scores competitively with Opus 4.8 on long-horizon coding tasks. Flat $3/$15 per 1M pricing keeps it accessible, but it still trails Fable 5 and GPT 5.6 Sol on precise-execution benchmarks, and weights remain hosted-only until July 27. Best for teams doing long-context agentic coding who can tolerate verbosity and don't need immediate self-hosting.

Input
$3.00 / 1M
Output
$15.00 / 1M
Context
1M tokens, $0.30/1M cached
Llama 4 Maverick logo

Llama 4 Maverick

Meta

Conditional

The leading open-source model for teams that need self-hosting, data sovereignty, or want to avoid API vendor lock-in.

Input
Free (self-hosted) or $0.20 / 1M tokens (hosted)
Output
Free (self-hosted) or $0.60 / 1M tokens (hosted)
Context
10M tokens
MAI-Code-1-Flash logo

MAI-Code-1-Flash

Microsoft

Conditional

Microsoft's first in-house coding model — a small, fast, aggressively cost-optimized sparse MoE at $0.75/$4.50 per 1M tokens. Delivers a 16-point lead over Claude Haiku 4.5 on SWE-Bench Pro while using up to 60% fewer tokens. It is purpose-built for high-volume Copilot workflows — autocompletions, refactors, scaffolds — not a frontier model. Best inside the Copilot ecosystem; API access outside it is still rolling out.

Input
$0.75 / 1M tokens
Output
$4.50 / 1M tokens
Context
256K context
Muse Spark 1.1 logo

Muse Spark 1.1

Meta

Conditional

Meta's first paid AI model and first salvo in the commercial coding market — $1.25/$4.25 per 1M tokens with 1M context, native subagent orchestration, MCP/tool support, and computer-use capability at roughly 4x under Anthropic/OpenAI pricing. Leads on tool-use benchmarks but trails Opus 4.8 on pure coding by 7 points on SWE-Bench Pro. Closed-weight, US-only public preview. Best for teams prioritizing tool-use orchestration at low cost.

Input
$1.25 / 1M tokens
Output
$4.25 / 1M tokens
Context
1M tokens
Mythos 5 (Claude Mythos 5) logo

Mythos 5 (Claude Mythos 5)

Anthropic

Conditional

The strongest cybersecurity model in the world (83.8% CyberGym) — built on Fable 5's weights with safety classifiers lifted. Access was suspended after the Jun 12 export directive; partially restored Jun 26 to ~100 designated US critical-infrastructure organizations (Project Glasswing). Verdict stays CONDITIONAL: capability is elite, but access is gated by federal designation. Not a public API — usable only if your organization is on the Commerce Department list.

Input
$10 / 1M tokens
Output
$50 / 1M tokens
Context
1M tokens (restricted access)
O

Ornith-1.0

DeepReinforce

Conditional

A strong MIT-licensed open-source agentic coding model family that achieves frontier-competitive benchmarks — beating Claude Opus 4.7 on both Terminal-Bench 2.1 and SWE-Bench Verified — but requires self-hosting with significant GPU resources and comes from a new lab with limited track record.

Input
Free (MIT license)
Output
Free (MIT license)
Context
256K tokens
S

Sakana Fugu

Sakana AI

Conditional

Multi-agent orchestrator that matches frontier models on key benchmarks by routing across a pool of expert agents — but opaque routing and platform dependency limit transparency.

Input
$5 / 1M tokens
Output
$30 / 1M tokens
Context
272K tokens (>272K: $10/$45 per 1M)