GLM-5.3-Flash now bills at $0.15/$0.50 after its 50% launch discount ended

Z.ai logoZ.aiImportantSeptember 14, 2026Pricing
What happened
Z.ai's 50 percent GLM-5.3-Flash launch promotion ended at 24:00 UTC+8 on 9 September 2026; Flash now bills at the $0.15/$0.50 per 1M token list rate.
Why it matters
Any Flash cost model built on the $0.075/$0.25 promo is half the real bill, and our own entry stored that promo as standing while calling GLM-5.3 pricing unpublished.
What to do
Re-forecast GLM-5.3-Flash workloads at $0.15/$0.50 per 1M tokens; full-size GLM-5.3 is unchanged at $1.40/$4.40.

GLM-5.3-Flash bills at $0.15 per 1M input tokens and $0.50 per 1M output. The $0.075/$0.25 rate quoted at launch was a 50 percent promotion that ended at 24:00 UTC+8 on 9 September 2026, and Z.ai's pricing table no longer carries it. Full-size GLM-5.3 is unchanged at $1.40/$4.40, and our verdict stays conditional: the Flash price moved, the trust question did not.

What happened

When we read Z.ai's official pricing table on 9 September 2026, it listed GLM-5.3-Flash at $0.075 input and $0.25 output per million tokens, a 50 percent discount on the $0.15/$0.50 list rate ending 9 September 2026, 24:00 UTC+8. Read again on 14 September, the table shows Flash at its list rate with no discount note.

ModelInput / 1MCached input / 1MOutput / 1M
GLM-5.3$1.40$0.26$4.40
GLM-5.2$1.40$0.26$4.40
GLM-5.1$1.40$0.26$4.40
GLM-5$1.00$0.20$3.20
GLM-5.3-Flash (list rate, current)$0.15$0.03$0.50
GLM-5.3-Flash (launch promo, ended 9 Sep 2026)$0.075not recorded$0.25

Why it matters

Any Flash cost model built on $0.075/$0.25 is half the real bill. That is arithmetic, not a benchmark argument.

The full-size rate is the other fact worth reading twice: GLM-5.3 costs exactly what GLM-5.2 and GLM-5.1 cost. Z.ai built 5.3 through post-training on the GLM-5.2 base and charged nothing extra for it.

We got both halves wrong

We are reporting this against our own entry. GLM-5.3 carried two defects:

  • The Flash promo was stored as a standing rate. Our pricing context read "GLM-5.3-Flash $0.075/$0.25" with no sign it was time-limited or that a higher list rate sat behind it. The entry's pricing fields now show the $0.15/$0.50 list rate.
  • The standing verdict said pricing was unpublished. The summary claimed "per-token pricing is unpublished" while the entity's own fields carried $1.40/$4.40, a rate live since at least 20 August 2026. That summary is corrected when this post publishes.

The two defects share one cause: prose written at launch, never re-read against the fields beside it.

The rating holds at conditional because GLM-5.3's headline coding and cyber scores, including Terminal Bench 3.0 and CyberGym, come from Z.ai's own team; the third-party runs so far (Artificial Analysis, Proximal's FrontierSWE) do not cover them. A published price answers a transparency question, not a trust one.

What changes for you

  • Re-forecast every GLM-5.3-Flash workload at $0.15 input and $0.50 output per 1M tokens. If your number came from a launch write-up, double it.
  • Budget full-size GLM-5.3 at $1.40 / $4.40 with $0.26 cached input. If you already priced GLM-5.2, nothing changes.
  • Do not standardize on GLM-5.3's headline benchmark scores until someone outside Z.ai reproduces them.

FAQ

Is GLM-5.3-Flash still $0.075 per 1M input tokens? No. That was a 50 percent launch promotion that ended at 24:00 UTC+8 on 9 September 2026. The list rate is $0.15 input and $0.50 output per 1M tokens.

Did full-size GLM-5.3 get more expensive? No. It is $1.40 input and $4.40 output per 1M tokens, the same as GLM-5.2. Only the Flash tier moved.

Why is the verdict still conditional? Because Z.ai's own team ran the headline coding and cyber benchmarks, and no pricing fact changes that.

conditionalprevious pick
conditionalnew pick

Rating unchanged: the headline coding and cyber benchmarks are still run by Z.ai's own team. The standing summary is corrected: it said per-token pricing was unpublished when Z.ai's table lists $1.40/$4.40, and the entity had stored the expired GLM-5.3-Flash launch promotion as if it were the list rate.

What to do

  1. 1 Re-forecast every GLM-5.3-Flash workload at $0.15 input and $0.50 output per 1M tokens, not the $0.075/$0.25 launch promotion that ended 9 September 2026.
  2. 2 Budget full-size GLM-5.3 at $1.40 / $4.40 per 1M tokens with $0.26 cached input; if you already priced GLM-5.2, nothing changes.
  3. 3 Do not standardize on GLM-5.3's headline benchmark scores until a third party reproduces them.

Affected tools & models

Never need to catch up again

The weekly delta — only verdict changes and act-now items. No digest filler.

By subscribing you agree to our Privacy Policy. Unsubscribe anytime.