Qwen3.8-Max API Pricing Ships at $2/$6 per 1M Tokens

Alibaba Cloud logoAlibaba CloudFYIAugust 11, 2026Models
What happened
Qwen3.8-Max now has published per-token API pricing at $2 input / $6 output per million tokens on QwenCloud, replacing the credit-only Token Plan preview.
Why it matters
At $2/$6 it undercuts Kimi K3 by 2.5x on output and matches Grok 4.5 at the same rate — the first Max-class Qwen model with transparent per-token costs.
What to do
Test Qwen3.8-Max on coding and agent workloads at the published $2/$6 rate; monitor Hugging Face and ModelScope for open weights (promised this week) before committing to self-hosted deployments.

Qwen3.8-Max, Alibaba's 2.4-trillion-parameter Mixture-of-Experts flagship, now has published per-token API pricing on QwenCloud: $2.00 per million input tokens and $6.00 per million output tokens. This replaces the credit-only Token Plan pricing that had been in place since the July 19 preview. Our verdict stays pending — pricing is now transparent, but open weights, independent benchmarks, and a disclosed license are still missing.

What happened

On August 3, 2026, Alibaba moved Qwen3.8-Max from preview to general availability, publishing standard pay-as-you-go per-token pricing at $2/$6 per million tokens (QwenCloud, 2026). For seven weeks prior, the model was gated behind QwenCloud's Token Plan subscription tiers — users bought monthly credit bundles without a standard rate card.

The model is a 2.4T-parameter MoE architecture activating roughly 95 billion parameters per token, with a 1-million-token context window and native multimodal input covering text, images, video, and documents (QwenCloud, 2026). It is accessible via OpenAI and Anthropic-compatible APIs, making it drop-in compatible with Cursor, Claude Code, Codex, and OpenCode.

Alibaba committed to releasing open weights for both Qwen3.8-Max and the smaller Qwen3.8-27B companion "within about a week" of the August 3 GA launch (Alibaba Qwen, 2026) — which would land the release in the week of August 10. As of today (August 11), no repository has appeared on Hugging Face or ModelScope, and no license has been disclosed.

Why it matters

At $2/$6, Qwen3.8-Max is aggressively priced for a frontier-scale model:

ModelInput (per 1M tokens)Output (per 1M tokens)
Qwen3.8-Max$2.00$6.00
Kimi K3$3.00$15.00
Grok 4.5$2.00$6.00
Claude Opus 5$5.00$25.00
GPT-5.6 Sol$5.00$30.00

It undercuts Kimi K3 by 2.5x on output tokens and matches Grok 4.5 at the same rate. Combined with the 1M-token context window and OpenAI/Anthropic API compatibility, developers can now test a frontier-scale model in existing toolchains at a known, competitive cost.

The open-weight promise is the real decision gate. If Alibaba ships the weights under a permissive license, Qwen3.8-Max becomes the largest open-weight model ever released — bigger in practical terms than Kimi K3, whose license imposes commercial restrictions and mandatory disclosures for Model-as-a-Service providers. If the weights don't ship, or ship under restrictive terms, the model remains one more API-only option.

Alibaba's self-reported benchmarks place Qwen3.8-Max at 86.1 on OSWorld-Verified, ahead of both GPT-5.6 Sol Max (83.2) and Claude Fable 5 (85.0) on that specific test (Developers Digest, 2026). But independent verification remains limited — as of August 8, neither Artificial Analysis nor LMArena had scored the model independently.

What changes for you

If you're already testing Qwen3.8-Max via Token Plan, you can now model your costs at $2/$6 before committing to production workloads. Pin the qwen3.8-max model identifier explicitly — a generic "latest" alias risks a future model swap silently changing your cost or behavior profile.

If you're waiting for open weights, the week of August 10 is the window Alibaba committed to. Monitor the Qwen organization on Hugging Face and ModelScope for the Qwen3.8-Max and Qwen3.8-27B repositories. Until the license is published, treat the open-weight release as a conditional promise, not a done deal.

For production deployments, wait for independent benchmarks from Artificial Analysis or equivalent before committing. Alibaba's self-reported numbers are directionally strong but unverified.

FAQ

Is Qwen3.8-Max open source? Not yet. Alibaba committed to releasing open weights for both the Max flagship and the 27B companion within a week of the August 3 GA launch. As of August 11, neither model has appeared on Hugging Face or ModelScope, and no license has been disclosed.

Can I use it with my existing tools? Yes. Qwen3.8-Max supports both OpenAI and Anthropic-compatible API protocols, so it works with Cursor, Claude Code, Codex, OpenCode, and Qwen Code out of the box. Use the qwen3.8-max model ID and your QwenCloud API key.

How does it compare to Kimi K3? On pricing, Qwen3.8-Max is 2.5x cheaper on output ($6 vs. $15 per million tokens). On benchmarks, Alibaba's self-reported scores are competitive but unverified. K3 has published weights under a restrictive license; Qwen3.8-Max's weights are still pending.

What to do

  1. 1 Test Qwen3.8-Max on coding and agent workloads at the published $2 input / $6 output per million tokens
  2. 2 Pin the explicit model identifier qwen3.8-max — do not use a generic 'latest' alias
  3. 3 Monitor Hugging Face and ModelScope for the promised open-weight release this week
  4. 4 Wait for independent benchmarks from Artificial Analysis or LMArena before production deployment

Affected tools & models

Never need to catch up again

The weekly delta — only verdict changes and act-now items. No digest filler.

By subscribing you agree to our Privacy Policy. Unsubscribe anytime.