Kimi K3 Tops Arena Coding at $3/$15, Open Weights July 27

Moonshot AI logoMoonshot AIImportantJuly 17, 2026Models
What happened
Moonshot AI launched Kimi K3, a 2.8T open-weight MoE model that beat Claude Fable 5 for #1 on Arena Code WebDev (1679) at $3/$15 per 1M tokens.
Why it matters
It's the third Chinese frontier-competitive open-weight model this month — Hy3, GLM-5.2, and now K3 — while Google's Gemini 3.5 Pro was confirmed 'months behind' the same day, making the open-weight frontier a genuine challenge to US closed models.
What to do
Watch the July 27 weight release, and if K3 ships open as promised, evaluate it for coding workloads where Fable 5's $10/$50 pricing is unsustainable — K3 at $3/$15 with 90%+ cache-hit rates sets a new cost baseline.

Kimi K3 is the largest open-weight model ever released — and its first act is beating Claude Fable 5 for the #1 spot on Arena Code WebDev at three times lower cost.

Moonshot AI launched the 2.8-trillion-parameter Mixture-of-Experts model on July 16, 2026, at a flat $3/$15 per million tokens (Moonshot AI, 2026). API access is live now on the Kimi app, Kimi Code, and through an OpenAI-compatible endpoint. Open weights are promised by July 27 — if delivered, K3 would be the largest downloadable model by a margin of nearly 3x over the next biggest open-weight release.

What Kimi K3 delivers

K3 scored 1679 on Arena Code WebDev — the top spot, ahead of both Fable 5 and GPT-5.6 Sol X-High on front-end coding tasks (LMSYS Arena, 2026). On Artificial Analysis's broader intelligence index, it lands at 57 — comparable to Claude Opus 4.8 and GPT-5.5, though still behind Fable 5 and Sol on general reasoning benchmarks.

Two new architectural components separate K3 from its predecessor, Kimi K2: Kimi Delta Attention (KDA), a hybrid linear attention mechanism, and Attention Residuals (AttnRes). Together they deliver roughly 2.5x better scaling efficiency. The 896-expert MoE activates only 16 experts per token, keeping inference costs low (Moonshot AI, 2026).

ModelInput PriceOutput PriceArena WebDevWeights
Kimi K3$3 / 1M$15 / 1M1679 (#1)Jul 27 (promised)
Claude Fable 5$10 / 1M$50 / 1M#2Closed
GPT-5.6 Sol$5 / 1M$30 / 1M#3Closed

Moonshot is pruning its older model line alongside the launch: K2.5 and the legacy moonshot-v1 series close to new users, with full sunset on August 31 (Moonshot AI, 2026).

Why it matters

This is the third Chinese open-weight model to enter frontier-competitive territory in 30 days. Tencent shipped Hy3 on July 6 with Apache 2.0 licensing and strong tool-calling. Z.ai launched GLM-5.2 on June 13 under MIT license with 1M-token context. Now K3 arrives with a #1 Arena WebDev score and the largest parameter count ever released openly.

The timing is brutal for Google. On the same day K3 launched, Bloomberg confirmed Gemini 3.5 Pro is "months behind" internal goals — the US model that was supposed to provide a third frontier option alongside Anthropic and OpenAI is still not generally available (Bloomberg, 2026).

Moonshot is also reportedly raising fresh capital at a $31.5 billion valuation — a substantial jump from its previous raise, though the $4.3B Series C figure from January 2026 is hard to independently verify (Reuters, 2026). That's the kind of trajectory usually reserved for labs shipping closed frontier models. The message is clear: open-weight models are attracting frontier-scale investment.

The caveat: K3 doesn't lead on general reasoning. Fable 5 and GPT-5.6 Sol still hold the edge on precise-execution and analytical benchmarks (Moonshot AI, 2026). This is a coding and cost-efficiency play, not a claim to the overall frontier crown. And the weights aren't downloadable yet — every prior Kimi flagship did ship open weights, but until July 27 it's a promise, not a fact.

What changes for you

If you're paying Fable 5's $10/$50 credit-only pricing for coding workloads, K3 resets the cost baseline. The model reports above 90% cache-hit rates on coding tasks, with cached input dropping to $0.30 per million tokens (Moonshot AI, 2026). For a team running heavy daily coding pipelines, the cost difference is not marginal — it's the difference between a sustainable workflow and a rationed one.

Open-weight access, if delivered on schedule, also means self-hosted deployment for teams with sufficient GPU infrastructure (64+ accelerators). No API dependency, no rate limits, no data leaving your infrastructure. We recommend evaluating K3 alongside GLM-5.2 and Hy3 once all three are downloadable.

Cache-hit pricing is the sleeper advantage here. At $0.30 per million tokens for cached input, repeated coding passes — re-reading the same codebase across iterations — become dramatically cheaper. K3's architecture was designed for this pattern, and the 90%+ hit rates Moonshot reports on coding workloads suggest it works as intended.

FAQ

Is Kimi K3 actually better than Fable 5?

On front-end coding and web development, yes — it leads Arena WebDev by a measurable margin (1679 vs Fable 5's score). On general reasoning, precise execution, and cybersecurity benchmarks, no — Fable 5 and GPT-5.6 Sol still hold the edge. Think of K3 as a coding and efficiency specialist, not a general-purpose frontier replacement.

When can I download the weights?

Moonshot promises open-weight release by July 27. Every prior Kimi flagship (K1.5, K2, K2.5) shipped weights, so the track record supports the claim. Until then, K3 is API-only through the Kimi platform and OpenAI-compatible endpoints.

How does K3 compare to GLM-5.2 and Hy3?

K3 is larger — 2.8T parameters vs Hy3's 295B (21B active) — and scores higher on Arena WebDev. Hy3 wins on permissive Apache 2.0 licensing and lower self-hosting requirements. GLM-5.2 has the advantage of MIT licensing and a longer track record. All three are open-weight and permissively licensed, making this 30-day window unprecedented in the open model landscape.

What to do

  1. 1 Evaluate Kimi K3 on your coding benchmarks once weights ship July 27 — at $3/$15 with 90%+ cache-hit rates, it resets the cost baseline for heavy coding workloads.
  2. 2 If self-hosting, budget for 64+ accelerators — K3's 2.8T MoE architecture demands serious GPU resources.

Affected tools & models

Never need to catch up again

The weekly delta — only verdict changes and act-now items. No digest filler.

By subscribing you agree to our Privacy Policy. Unsubscribe anytime.