AI Models

54

Find the right model for your task. Filter by strength, compare pricing, and decide.

Filter by provider

54models
Astra logo

Astra

OpenAI

pending

Astra is OpenAI's first model to trigger the 'Critical' cybersecurity threshold under its Preparedness Framework — internal evaluations found it may autonomously exploit zero-day vulnerabilities and solved 10 open math problems for ~$2,000. But Astra is not a shipping product: development is actively paused as of August 2026 for security hardening, with no release timeline or pricing. Watch for the safety-governance outcome — this is the first time a major lab has publicly classified its own model as critically dangerous.

Input
Not yet published
Output
Not yet published
Context
Development slowed Aug 2026. No GA date.
Claude Fable 5 logo

Claude Fable 5

Anthropic

Conditional

Anthropic's Mythos-class frontier model — the strongest publicly-available Claude yet, leading every major benchmark at launch. The verdict stays CONDITIONAL on two live constraints: credit-only pricing at $10/$50 per 1M tokens with no subscription path, and a safety classifier that over-triggers on routine coding tasks, silently routing them to an Opus 4.8 fallback. Reach for Fable 5 when the capability delta justifies the cost; for routine work, Sonnet 5 is the more predictable alternative.

Input
$10 / 1M tokens
Output
$50 / 1M tokens
Context
1M context, 128K max output
Claude Haiku 4.5 logo

Claude Haiku 4.5

Anthropic

Recommended

Anthropic's value baseline at $1/$5 per 1M tokens — the fast, low-cost tier rivals benchmark against, with 200K context and 64K output. Best for high-volume latency-sensitive work: classification, extraction, summarization, routing. Reach for Opus 4.8 or Sonnet 5 when deep multi-step reasoning is needed.

Input
$1 / 1M tokens
Output
$5 / 1M tokens
Context
200K tokens
Claude Opus 4.8 logo

Claude Opus 4.8

Anthropic

Recommended

Anthropic's flagship reasoning model and the practical frontier incumbent for deployable capability — 1M-token context at $5/$25 per 1M tokens, included across all subscription plans. Our pick for review-heavy work: code review, constrained refactoring, and deliberate-paced tasks. Strong agentic coding, top-tier analytical reasoning, and stable subscription access make it the default flagship for most production teams.

Input
$5 / 1M tokens
Output
$25 / 1M tokens
Context
1M tokens
Claude Opus 5 logo

Claude Opus 5

Anthropic

Recommended

Near-Fable 5 intelligence at half the price ($5/$25 per 1M tokens). SOTA on Frontier-Bench and CursorBench with 85% fewer cyber safeguards than Fable 5. The workhorse frontier model for everyday coding and knowledge work.

Input
$5.00 / 1M tokens
Output
$25.00 / 1M tokens
Context
200K tokens
Claude Sonnet 4.6 logo

Claude Sonnet 4.6

Anthropic

Conditional

Superseded by Claude Sonnet 5 (shipped Jun 30, 2026) as the default mid-tier workhorse. Sonnet 4.6's $3/$15 pricing and 1M-token context are now matched by a successor with a 13.4-point Terminal-Bench 2.1 jump at the same rate card. Choose 4.6 only for pinned deployments or where Sonnet 5's ~30% higher token count and removed temperature/top_p/top_k are disqualifying.

Input
$3 / 1M tokens
Output
$15 / 1M tokens
Context
1M tokens
Claude Sonnet 5 logo

Claude Sonnet 5

Anthropic

Recommended

Anthropic's most agentic Sonnet delivers near-Opus 4.8 performance at roughly half the price, with Terminal-Bench 2.1 +13.4 points over Sonnet 4.6. The new tokenizer inflates counts ~30%, so intro pricing ($2/$10 per MTok through Aug 31) keeps migration cost-neutral, but teams should recount budgets before switching. It replaces Sonnet 4.6 as the default across all Claude plans, ideal for agentic coding, brownfield debugging, and production pipelines — not for the hardest frontier reasoning or cybersecurity tasks.

Input
$3 / 1M
Output
$15 / 1M
Context
1M tokens
D

DeepSeek V4 Flash

DeepSeek

Recommended

DeepSeek's volume-tier open-weight model at 284B parameters (13B active MoE) with a 1M-token context window. SWE-bench Verified 79.0% at $0.22 input and $0.66 output per 1M tokens off-peak, doubling during the seven daily peak hours. The August 2026 price rise ended its run as the cheapest credible API (GPT-5.6 Luna now undercuts it on input at $0.20), but it remains a strong value pick, and the MIT license means self-hosting sidesteps API pricing entirely. Still the default workhorse for high-throughput agent pipelines that can run off-peak.

Input
$0.22 / 1M tokens
Output
$0.66 / 1M tokens
Context
1M tokens. Peak 2x: 01-04 + 06-10 UTC
D

DeepSeek V4 Pro

DeepSeek

Recommended

The strongest open-weight coding model we've tested: SWE-bench Verified 80.6%, LiveCodeBench 93.5% and 3206 Codeforces Elo. At $0.66 input and $1.98 output per 1M tokens off-peak it stays far below Western flagship pricing even after the August 2026 rise, though rates double during the seven daily peak hours. The default for self-hosted agentic coding.

Input
$0.66 / 1M tokens
Output
$1.98 / 1M tokens
Context
1M tokens. Peak 2x: 01-04 + 06-10 UTC
F

FLUX 3

Black Forest Labs

Conditional

FLUX 3 is Black Forest Labs' first unified multimodal foundation model — a single architecture jointly trained on image, video, audio, and action prediction. It generates 20-second video clips with native synchronized audio, beating Runway Gen-4.5 by 77% and Luma Ray 3.2 by 93% in vendor-reported preference tests, and its FLUX-mimic variant is already deployed on Audi production lines for soft-body manipulation. But three of four components remain unreleased (only Video and Action are in gated early access), no pricing has been published, no independent benchmarks exist, and open-weight FLUX 3 Dev won't ship until later 2026. Conditional: groundbreaking architecture with real-world robotics validation, but not yet a product you can buy or independently evaluate.

Input
Not yet published
Output
Not yet published
Context
Gated early access
G

GLM-5.2

Z.ai

Conditional

Top open-weight coding model with 1M-token context and MIT license — beats GPT-5.5 on multiple benchmarks at a fraction of cost, but still trails Claude Opus 4.8 on the hardest long-horizon tasks.

Input
$1.40 / 1M tokens
Output
$4.40 / 1M tokens
Context
1M tokens
G

GLM-5.3

Z.ai

Conditional

GLM-5.3 is Z.ai's latest open-weights coding and cyber-defense model — the same base as GLM-5.2, every gain from post-training. But the headline numbers are still vendor-run, the open weights are held ~2 weeks over emergent security capability, and per-token pricing is unpublished. Reach for defensive security and agentic automation now; wait for the weights and independent audits before standardizing.

Input
N/A (rate not yet published)
Output
N/A (rate not yet published)
Context
1M context
GPT-4.1 logo

GPT-4.1

OpenAI

Recommended

The default migration target for GPT-4 Turbo workloads: same request shape, lower cost, 1M-token context — convert legacy function_call payloads to the tools format and move on.

Input
$2 / 1M tokens
Output
$8 / 1M tokens
Context
1M tokens
GPT-4o (Retired from ChatGPT, API-only) logo

GPT-4o (Retired from ChatGPT, API-only)

OpenAI

Conditional

GPT-4o (May 2024) has been superseded by the GPT-5.x family and delisted from OpenAI's flagship API pricing page. It remains available through the API but is no longer a current-generation model. Use only where migration to GPT-5.x is blocked; otherwise, GPT-5.5 or GPT-5.4 offer better performance at comparable cost.

Input
$2.50 / 1M tokens
Output
$10 / 1M tokens
Context
128K tokens
GPT-5.5 logo

GPT-5.5

OpenAI

Recommended

OpenAI's general-purpose commercial frontier model and the deployable OpenAI anchor while the GPT-5.6 family stays government-gated — $5/$30 per 1M tokens with 1M+ context. Strong reasoning, broad ecosystem support, and full public availability make it a solid production default, though cheaper models now match it on specific tasks. For most teams, it remains the baseline newer models are measured against.

Input
$5 / 1M tokens
Output
$30 / 1M tokens
Context
1M+ tokens (922K in / 128K out)
GPT-5.5-Cyber logo

GPT-5.5-Cyber

OpenAI

Conditional

85.6% CyberGym — top single-model score — but restricted-access for vetted defenders only.

Input
N/A
Output
N/A
Context
Trusted Access for Cyber; not commercial API
GPT-5.6 Luna logo

GPT-5.6 Luna

OpenAI

Conditional

GPT-5.6 Luna is OpenAI's budget tier — near-GPT-5.5 capability at $1/$6 per 1M tokens, the strongest capability-per-dollar value in the GPT-5.6 family. Now generally available via API and Codex after the July 9 government-gate lift. Remains conditional: all benchmarks are vendor-reported, Luna lacks activation classifiers, and it is not in the standard ChatGPT picker.

Input
$0.20 / 1M tokens
Output
$1.20 / 1M tokens
Context
Fastest tier. 80% cut Jul 30 → $0.20/$1.20.
GPT-5.6 Sol logo

GPT-5.6 Sol

OpenAI

Recommended

GPT-5.6 Sol is the most cost-effective frontier model, leading on coding benchmarks and competitive on hard reasoning at a fraction of the cost of peers. Verdict reinforced by GPT-Red safety-hardening disclosure — the model was adversarially trained against a dedicated self-play RL red-teaming model at frontier scale, making it 6x more robust to prompt injections.

Input
$5 / 1M tokens
Output
$30 / 1M tokens
Context
Same $5/$30. Fast mode: 2.5× at 2× cost.
GPT-5.6 Terra logo

GPT-5.6 Terra

OpenAI

Conditional

GPT-5.6 Terra is OpenAI's mid-tier workhorse — GPT-5.5-class capability at half the cost ($2.50/$15 per 1M tokens). Now generally available via API and Codex after the July 9 government-gate lift. Remains conditional: all benchmarks are vendor-reported without independent reproduction, and Terra is not selectable in the standard ChatGPT model picker.

Input
$2.00 / 1M tokens
Output
$12.00 / 1M tokens
Context
Mid-tier. 20% cut Jul 30 → $2.00/$12.00.
GPT-5.6-Cyber logo

GPT-5.6-Cyber

OpenAI

Conditional

GPT-5.6-Cyber is OpenAI's purpose-trained cybersecurity model, built on GPT-5.6 Sol and available exclusively through the gated Daybreak Red program. It is the first model to reach OpenAI's 'High' cyber capability threshold under its Preparedness Framework, with exploit-chain and zero-day performance that far exceeds general-purpose models. Reach for it if you are an authorized red team or vulnerability researcher; everyone else should use GPT-5.6 Sol via Daybreak Blue.

Input
$12.50 / 1M tokens
Output
$75 / 1M tokens
Context
400K tokens
GPT-Live-1 logo

GPT-Live-1

OpenAI

Conditional

OpenAI's first full-duplex voice model — listens and speaks simultaneously, a leap beyond turn-based Advanced Voice Mode. Delegates complex reasoning to GPT-5.5 in the background while maintaining conversational flow. Verdict CONDITIONAL: it's the best ChatGPT voice experience available, but it's a product-integrated model with no API and uneven language support.

Input
ChatGPT subscription
Output
ChatGPT subscription
Context
ChatGPT Voice only
GPT-Live-1 mini logo

GPT-Live-1 mini

OpenAI

Conditional

The free-tier variant of GPT-Live-1 — same full-duplex architecture in a smaller, faster package, serving as the default voice model for ChatGPT Free users. Shares the core innovations: continuous interaction, wake word support, and GPT-5.5 delegation. Same limitations apply: no API, uneven language quality. A substantial upgrade over the old turn-based voice mode for free users.

Input
Free (ChatGPT)
Output
Free (ChatGPT)
Context
ChatGPT Voice only
GPT-Live-Transcribe logo

GPT-Live-Transcribe

OpenAI

Recommended

OpenAI's low-latency streaming speech-to-text model delivers live transcript deltas at $0.017/min with tunable latency and context-aware transcription. It improves on gpt-realtime-whisper's error rate and supports real-time keyword hints, language hints, and mid-session configuration updates via WebSocket or WebRTC. The tradeoff: ~3.8x the cost of batch GPT-Transcribe, and no diarization, word timestamps, or confidence scores. Best for live captioning, meeting assistants, and voice agents where sub-second transcript display matters more than per-minute cost.

Input
$0.017 / minute
Output
N/A
Context
Per-minute realtime audio billing
GPT-Transcribe logo

GPT-Transcribe

OpenAI

Recommended

OpenAI's latest batch speech-to-text model scores 3.31% AA-WER at $0.0045/min — 25% cheaper and measurably more accurate than its gpt-4o-transcribe predecessor. It supports context prompting, keyword hints, and language detection across 22+ languages, making it the best default for asynchronous file transcription on the OpenAI platform. Caveat: no diarization or word-level timestamps; use gpt-4o-transcribe-diarize for speaker labels or gpt-live-transcribe for real-time needs. Best for developers building meeting transcription, media processing, and voice interfaces on the OpenAI API.

Input
$0.0045 / minute
Output
N/A
Context
Per-minute audio billing
Gemini 2.5 Pro logo

Gemini 2.5 Pro

Google

Recommended

Now the budget long-context pick: a 40% price cut (Jun 2026) made full-corpus passes affordable at mid-size budgets — the longest context window available with strong Google ecosystem integration.

Input
$1.25 / 1M tokens
Output
$10 / 1M tokens
Context
1M tokens
Gemini 3.1 Pro logo

Gemini 3.1 Pro

Google

Conditional

Gemini 3.1 Pro is Google's most capable publicly available Pro-tier model, delivering strong reasoning with a 1M-token context window at competitive pricing — and it's the migration target for GitHub Copilot users leaving deprecated Gemini 2.5 Pro today. But it's decidedly previous-gen: Gemini 3.6 and 3.5 Flash now beat it on coding and agentic benchmarks at lower cost. Stick with it if you're deep in the Google Cloud / Vertex AI ecosystem; otherwise, the Flash line or competitors deliver more for less.

Input
$2 / 1M
Output
$12 / 1M
Context
1M tokens
Gemini 3.5 Flash logo

Gemini 3.5 Flash

Google

Recommended

Google's cost-efficiency standout for agentic desktop automation at $1.50/$9 per 1M tokens with ~1M context. Native Computer Use scores 78.4 on OSWorld-Verified — within 0.3 points of GPT-5.5 at roughly one-third the cost. THE relevant caveat: the Computer Use tool is a public preview not yet in the production API, so verify endpoint availability before wiring in.

Input
$1.50 / 1M tokens
Output
$9 / 1M tokens
Context
1M tokens
Gemini 3.5 Flash Cyber logo

Gemini 3.5 Flash Cyber

Google

Caution

Gemini 3.5 Flash Cyber is Google's first cybersecurity-specialized model, fine-tuned from 3.5 Flash to find, validate, and patch vulnerabilities. Inside CodeMender's multi-agent architecture, it reaches competitive performance on CyberGym against much larger models like Mythos at a fraction of the cost. But the restrictions are severe: it runs exclusively within CodeMender, available only to governments and trusted partners via a limited-access pilot with no public API, no published benchmark numbers, and no timeline for broader release. For the vast majority of developers, this model is simply inaccessible — watch for access expansion.

Input
N/A
Output
N/A
Context
CodeMender-only, government-gated pilot
Gemini 3.5 Flash-Lite logo

Gemini 3.5 Flash-Lite

Google

Recommended

Gemini 3.5 Flash-Lite is Google's fastest, cheapest 3.5-series model at 350 tok/s and $0.30/$2.50. The generational leap from 3.1 Flash-Lite is the real story — it dramatically outperforms its predecessor on key agentic benchmarks and beats the full-size Gemini 3 Flash on coding and computer-use tasks. For developers scaling high-throughput agentic subagents, search pipelines, and document processing, Flash-Lite's price-to-performance ratio is unmatched in Google's lineup.

Input
$0.30 / 1M
Output
$2.50 / 1M
Context
1M context, 64K output, 350 tok/s
Gemini 3.5 Pro logo

Gemini 3.5 Pro

Google

pending

Google's pre-GA frontier model — now on its fourth postponement and 'months behind schedule' per Bloomberg. The 2M-token context, native multimodal pipeline, and Deep Think reasoning remain real capabilities, but no confirmed GA timeline exists. Coding capability is the core shortfall.

Input
Not yet published
Output
Not yet published
Context
Est. $5-15/$30-60 per 1M. 2M ctx. Not GA.
Gemini 3.6 Flash logo

Gemini 3.6 Flash

Google

Recommended

Gemini 3.6 Flash is Google's new default Flash-tier workhorse — it's both better and cheaper than 3.5 Flash, scoring higher on every benchmark while cutting the output price from $9.00 to $7.50 per million tokens. The 17% token-efficiency gain means lower per-task cost even before the price cut, and built-in Computer Use via the API makes it a credible agentic coding option. For Google-ecosystem developers, this is the clear upgrade — there is no reason to stay on 3.5 Flash.

Input
$1.50 / 1M
Output
$7.50 / 1M
Context
1M context, 64K output
Gemini 3.7 Flash logo

Gemini 3.7 Flash

Google

Recommended

Gemini 3.7 Flash is Google's strongest Flash workhorse yet for coding and agents — a genuine step up from 3.6 Flash at half the prior price through 2026. It still trails the frontier on the hardest terminal and computer-use tasks. A clear upgrade for cost-sensitive and Google-ecosystem agent builders; not a frontier-code replacement.

Input
$0.75 / 1M
Output
$3.75 / 1M
Context
1M context, 64K output
Gemini 3.8 Flash logo

Gemini 3.8 Flash

Google

Recommended

Gemini 3.8 Flash is Google's strongest Flash workhorse yet — a genuine capability upgrade over the recommended 3.7 Flash at the same introductory price, with real gains across software engineering, agents and professional domains. The deciding caveat is Google's own: 3.8 "works harder," so per-task token burn runs well above 3.7, which remains the efficiency-first pick. Adopt 3.8 for long-horizon coding, agents and finance or legal work — benchmark your own workload before the intro price expires.

Input
$0.75 / 1M
Output
$3.75 / 1M
Context
1M context, 64K output
Gemini 4 logo

Gemini 4

Google

pending

Gemini 4 is Google DeepMind's confirmed next-generation frontier model — pre-training began July 21, 2026 on what Google calls its 'most ambitious' run yet. Sundar Pichai has positioned it as the answer to coding and agentic gaps, targeting where the frontier will be at launch, not where it sits today. But everything else is unknown: no benchmarks, no pricing, no specs, no GA release date, and the delayed Gemini 3.5 Pro's relationship to it remains unresolved. Watchlist-only until Google ships something concrete.

Input
Not yet published
Output
Not yet published
Context
Pre-training started Jul 2026. No GA date.
G

Grok 4.5

SpaceXAI

Conditional

Competitively priced coding/agentic model at $2/$6 per 1M tokens with 500K context — xAI's published benchmarks place it mid-pack behind Fable 5 and GPT-5.5, though it leads Opus 4.8 on DeepSWE and Terminal-Bench 2.1. The real story is cost efficiency: 4.2x more token-efficient than Opus 4.8 on SWE-Bench Pro. Best for teams needing strong coding at a fraction of frontier pricing. Not yet available in the EU.

Input
$2 / 1M tokens
Output
$6 / 1M tokens
Context
500K tokens
G

Grok 4.6

SpaceXAI

Conditional

Frontier-competitive at $2/$6 with a 61 AA Intelligence Index (ties GPT-5.6 Sol) and GDPVal-AA behind only Claude Opus 5 — strong out of the gate, but one evaluator deep and one day old, so hold production bets until broader benchmarks land.

Input
$2.00 / 1M tokens
Output
$6.00 / 1M tokens
Context
Delayed past Aug 7. Now ~Aug 10-14. 500K ctx.
G

Grok Voice Think Fast 2.0

SpaceXAI

Conditional

Grok Voice Think Fast 2.0 is SpaceXAI's next-generation speech-to-speech model with a parallel reasoning architecture that makes it substantially smarter than conventional voice models with no latency penalty. Its transcription accuracy destroys dedicated STT competitors in noisy environments — the deciding factor for production voice agents. But at $0.08/min, developer/API-only with no consumer-facing assistant, it's a premium specialist — not for cost-sensitive or general-purpose voice use.

Input
$0.08 / min of audio (speech-to-speech)
Output
Included in speech-to-speech rate
Context
API-only
H

Hy3

Tencent

Conditional

Apache 2.0 open-weight MoE agentic model (295B total, 21B active) with strong tool-calling reliability and anti-hallucination behavior, but coding trails GLM-5.2 and self-hosting demands 8+ GPUs. Free OpenRouter tier through Jul 21 — afterward $0.20/$0.80 per 1M tokens. Best for self-hosted coding agents and long-context document work where permissive licensing matters more than absolute benchmark scores.

Input
$0.20 / 1M
Output
$0.80 / 1M
Context
256K tokens
I

Inkling

Thinking Machines Lab

Conditional

Inkling is the strongest reason yet to take open-weight multimodal AI seriously. It delivers competitive reasoning and coding with best-in-class open-weight safety and controllable thinking effort, all under Apache 2.0. But Thinking Machines is honest that it's not the strongest overall — factuality lags (SimpleQA 43.9%) and self-hosting demands serious hardware (2TB VRAM). Conditional: a compelling base for organizations that need to own and fine-tune a multimodal model, not for teams wanting peak off-the-shelf performance.

Input
$1.87 / 1M tokens
Output
$4.68 / 1M tokens
Context
Up to 1M tokens native
K

Kimi K3

Moonshot AI

Conditional

Strong model now under active White House investigation for industrial-scale IP theft. US Treasury threatens sanctions — legal and operational risk is material. Technical benchmarks unchanged.

Input
$3.00 / 1M
Output
$15.00 / 1M
Context
1M ctx, $0.30 cached. Weights live (Jul 27).
L

Laguna S 2.1

poolside

Conditional

Laguna S 2.1 is the first credible Western open-weight coding model in 11 months — a 118B MoE that activates only 8B parameters per token yet scores 70.2% on Terminal-Bench 2.1, beating models 10x its size. But it's not at the frontier — GPT-5.6 Sol and Kimi K3 sit far ahead — and known limitations include harness overfitting, JSON mangling in nested tool arguments, and no configurable thinking-effort control. Poolside's collapsed $14B Series C raises sustainability questions, making this best for self-hosted agentic coding where data sovereignty matters more than absolute benchmark leadership.

Input
$0.10 / 1M
Output
$0.20 / 1M
Context
OpenMDW-1.1 license, 1M context
Llama 4 Maverick logo

Llama 4 Maverick

Meta

Conditional

The leading open-source model for teams that need self-hosting, data sovereignty, or want to avoid API vendor lock-in.

Input
Free (self-hosted) or $0.20 / 1M tokens (hosted)
Output
Free (self-hosted) or $0.60 / 1M tokens (hosted)
Context
10M tokens
MAI-Code-1-Flash logo

MAI-Code-1-Flash

Microsoft

Conditional

Microsoft's first in-house coding model — a small, fast, aggressively cost-optimized sparse MoE at $0.75/$4.50 per 1M tokens. Delivers a 16-point lead over Claude Haiku 4.5 on SWE-Bench Pro while using up to 60% fewer tokens. It is purpose-built for high-volume Copilot workflows — autocompletions, refactors, scaffolds — not a frontier model. Best inside the Copilot ecosystem; API access outside it is still rolling out.

Input
$0.75 / 1M tokens
Output
$4.50 / 1M tokens
Context
256K context
MAI-Code-1.1-Flash logo

MAI-Code-1.1-Flash

Microsoft

Conditional

MAI-Code-1.1-Flash is Microsoft's refreshed small-tier Copilot coding model — a quarter the cost and 22% better on Terminal-Bench 2.1 than its June predecessor, now with native vision. It's built for high-volume iterative work rather than frontier reasoning, making it the cheapest serious coding model in Copilot. Reach for it on everyday completions, refactors, and image-assisted debugging; use a flagship for complex architecture.

Input
$0.20 / 1M
Output
$1.20 / 1M
Context
256K context
MAI-Cyber-1-Flash logo

MAI-Cyber-1-Flash

Microsoft

Conditional

Microsoft's first in-house cybersecurity model, built on the MAI-Thinking-1 family and embedded inside the MDASH multi-agent harness. MAI-Cyber-1-Flash handles ~90% of security workloads at roughly half the cost of prior configurations, reserving GPT-5.4 for the hardest 10%. It scores 95.95% on CyberGym — 12 points above Mythos 5 — but this is a vendor-reported benchmark from a tuned full-system configuration (harness + dual models), not an independent model-versus-model test. Available only through Azure AI Foundry with vetted enterprise access; no standalone pricing or public API.

Input
N/A (MDASH SCU consumption)
Output
N/A (MDASH SCU consumption)
Context
Azure AI Foundry vetted access
Muse Glimmer logo

Muse Glimmer

Meta

Conditional

Muse Glimmer is Meta's 30B open-weight agentic model, distilled from the closed Muse Spark frontier model and released under Apache 2.0. With 4-bit quantization it runs multi-step tool calling, multimodal reasoning, and failure recovery on a single 24GB consumer GPU without cloud dependencies. It leads its size class on agentic orchestration but trails Qwen3.6-27B on terminal and OS automation — reach for it if you need a capable local agent that stays offline, not if you need independent reproducibility or low-level OS control.

Input
Free (open weights)
Output
Free (open weights)
Context
131K tokens
Muse Spark 1.1 logo

Muse Spark 1.1

Meta

Conditional

Meta's first paid AI model and first salvo in the commercial coding market — $1.25/$4.25 per 1M tokens with 1M context, native subagent orchestration, MCP/tool support, and computer-use capability at roughly 4x under Anthropic/OpenAI pricing. Leads on tool-use benchmarks but trails Opus 4.8 on pure coding by 7 points on SWE-Bench Pro. Closed-weight, US-only public preview. Best for teams prioritizing tool-use orchestration at low cost.

Input
$1.25 / 1M tokens
Output
$4.25 / 1M tokens
Context
1M tokens
Muse Spark 1.2 logo

Muse Spark 1.2

Meta

Conditional

Meta's coding-focused model update, co-trained with the Muse Code harness — 82.9% on Terminal-Bench 2.1 and 59.3% on DeepSWE 1.1, second only to Claude Opus 5 on the former and trailing both Opus 5 and GPT-5.6 Terra on the latter. Standard pricing unchanged from 1.1 at $1.25 input / $4.25 output per 1M tokens; contributor tier drops to $0.10/$0.20 in exchange for training-data rights. All published benchmarks remain vendor-run with no independent reproduction. Best for cost-sensitive coding teams willing to test a model explicitly optimized for its own harness.

Input
$1.25 / 1M tokens
Output
$4.25 / 1M tokens
Context
1M tokens
Mythos 5 (Claude Mythos 5) logo

Mythos 5 (Claude Mythos 5)

Anthropic

Conditional

The strongest cybersecurity model in the world (83.8% CyberGym) — built on Fable 5's weights with safety classifiers lifted. Access was suspended after the Jun 12 export directive; partially restored Jun 26 to ~100 designated US critical-infrastructure organizations (Project Glasswing). Verdict stays CONDITIONAL: capability is elite, but access is gated by federal designation. Not a public API — usable only if your organization is on the Commerce Department list.

Input
$10 / 1M tokens
Output
$50 / 1M tokens
Context
1M tokens (restricted access)
O

Ornith-1.0

DeepReinforce

Conditional

A strong MIT-licensed open-source agentic coding model family that achieves frontier-competitive benchmarks — beating Claude Opus 4.7 on both Terminal-Bench 2.1 and SWE-Bench Verified — but requires self-hosting with significant GPU resources and comes from a new lab with limited track record.

Input
Free (MIT license)
Output
Free (MIT license)
Context
256K tokens
P

Palmyra X6

Writer

Conditional

Palmyra X6 is Writer's flagship agentic model for marketing and revenue workflows — post-trained on Z.ai's open-weight GLM-5.2 at $2/$8 per 1M tokens, a fraction of the frontier rate. Every headline number is Writer-run, and the GLM-5.2 provenance is a real procurement question. Adopt it inside Writer for cost-sensitive GTM agent work; look elsewhere for independently benchmarked general capability.

Input
$2 / 1M
Output
$8 / 1M
Context
1M context, 8K max output
Q

Qwen3.8-27B

Alibaba

Conditional

Qwen3.8-27B is Alibaba's 27-billion-parameter open-weight model, now shipped under Apache 2.0 with 262K native context and single-GPU deployment. It is the self-hostable companion to the 2.4T Qwen3.8-Max flagship. Benchmarks remain Alibaba-reported, so it lands at conditional — a real, runnable model whose vendor-claimed coding gains still await independent reproduction.

Input
Free (open weights — self-hosted)
Output
Free (open weights — self-hosted)
Context
TBD — context window not published
Q

Qwen3.8-Max

Alibaba

pending

Qwen3.8-Max is Alibaba's 2.4-trillion-parameter flagship, previewed July 2026 as a multimodal MoE model with a 984K context window. It is not yet GA — no per-token API pricing, no benchmark table, no open weights, and no model card have been published, and the "second only to Fable 5" claim is Alibaba's own unverified self-assessment. The preview is worth testing for coding and long-context agent workloads at $6/month entry cost, but wait for published benchmarks, per-token pricing, and the promised open-weight release before committing.

Input
$2.00 / 1M tokens
Output
$6.00 / 1M tokens
Context
984K tokens. Open weights promised week of Aug 10.
S

Sakana Fugu

Sakana AI

Conditional

Multi-agent orchestrator that matches frontier models on key benchmarks by routing across a pool of expert agents — but opaque routing and platform dependency limit transparency.

Input
$5 / 1M tokens
Output
$30 / 1M tokens
Context
272K tokens (>272K: $10/$45 per 1M)