K

Kimi K3

Moonshot AI · Released Jul 2026

Conditional

Moonshot AI's 2.8T MoE flagship — the largest open-weight model ever released — leads Arena WebDev front-end coding and scores competitively with Opus 4.8 on long-horizon coding tasks. Flat $3/$15 per 1M pricing keeps it accessible, but it still trails Fable 5 and GPT 5.6 Sol on precise-execution benchmarks, and weights remain hosted-only until July 27. Best for teams doing long-context agentic coding who can tolerate verbosity and don't need immediate self-hosting.

Is it right for you?

Good for

  • Front-end and 3D work with image feedback — Arena WebDev #1 (1679, preliminary)
  • Long-horizon autonomous coding — MiniTriton compiler, 48-hour chip design demonstrations
  • Multi-agent orchestration — K3 Swarm Max variant for large-scale parallel processing
  • Large-context knowledge work — 1M window for whole-repo and multi-document research

Not good for

  • Precise-execution tasks — still trails Fable 5 and GPT 5.6 Sol
  • Self-hosted deployment — requires 64+ accelerators, weights not yet downloadable
  • Cost-sensitive workloads — AA measured high output verbosity driving up task costs

Pricing

Input

$3.00 / 1M

Output

$15.00 / 1M

Context

1M tokens, $0.30/1M cached

View full pricing

Benchmarks

BenchmarkScoreSource
DeepSWE67.5 Source
Terminal-Bench 2.188.3 Source
FrontierSWE81.2 Source
Program Bench77.8 Source
SWE Marathon42.0 Source
Arena WebDev1679 (#1, preliminary) Source
AA Intelligence Index57 Source

No verdict changes yet

The clock starts day one — changes land here as our verdict evolves.

Verification log

  • Pricing— No changes

    Automated agent

How we evaluate