GLM-5.3-Flash now bills at $0.15/$0.50 after its 50% launch discount ended
- What happened
- Z.ai's 50 percent GLM-5.3-Flash launch promotion ended at 24:00 UTC+8 on 9 September 2026; Flash now bills at the $0.15/$0.50 per 1M token list rate.
- Why it matters
- Any Flash cost model built on the $0.075/$0.25 promo is half the real bill, and our own entry stored that promo as standing while calling GLM-5.3 pricing unpublished.
- What to do
- Re-forecast GLM-5.3-Flash workloads at $0.15/$0.50 per 1M tokens; full-size GLM-5.3 is unchanged at $1.40/$4.40.
GLM-5.3-Flash bills at $0.15 per 1M input tokens and $0.50 per 1M output. The $0.075/$0.25 rate quoted at launch was a 50 percent promotion that ended at 24:00 UTC+8 on 9 September 2026, and Z.ai's pricing table no longer carries it. Full-size GLM-5.3 is unchanged at $1.40/$4.40, and our verdict stays conditional: the Flash price moved, the trust question did not.
What happened
When we read Z.ai's official pricing table on 9 September 2026, it listed GLM-5.3-Flash at $0.075 input and $0.25 output per million tokens, a 50 percent discount on the $0.15/$0.50 list rate ending 9 September 2026, 24:00 UTC+8. Read again on 14 September, the table shows Flash at its list rate with no discount note.
| Model | Input / 1M | Cached input / 1M | Output / 1M |
|---|---|---|---|
| GLM-5.3 | $1.40 | $0.26 | $4.40 |
| GLM-5.2 | $1.40 | $0.26 | $4.40 |
| GLM-5.1 | $1.40 | $0.26 | $4.40 |
| GLM-5 | $1.00 | $0.20 | $3.20 |
| GLM-5.3-Flash (list rate, current) | $0.15 | $0.03 | $0.50 |
| GLM-5.3-Flash (launch promo, ended 9 Sep 2026) | $0.075 | not recorded | $0.25 |
Why it matters
Any Flash cost model built on $0.075/$0.25 is half the real bill. That is arithmetic, not a benchmark argument.
The full-size rate is the other fact worth reading twice: GLM-5.3 costs exactly what GLM-5.2 and GLM-5.1 cost. Z.ai built 5.3 through post-training on the GLM-5.2 base and charged nothing extra for it.
We got both halves wrong
We are reporting this against our own entry. GLM-5.3 carried two defects:
- The Flash promo was stored as a standing rate. Our pricing context read "GLM-5.3-Flash $0.075/$0.25" with no sign it was time-limited or that a higher list rate sat behind it. The entry's pricing fields now show the $0.15/$0.50 list rate.
- The standing verdict said pricing was unpublished. The summary claimed "per-token pricing is unpublished" while the entity's own fields carried $1.40/$4.40, a rate live since at least 20 August 2026. That summary is corrected when this post publishes.
The two defects share one cause: prose written at launch, never re-read against the fields beside it.
The rating holds at conditional because GLM-5.3's headline coding and cyber scores, including Terminal Bench 3.0 and CyberGym, come from Z.ai's own team; the third-party runs so far (Artificial Analysis, Proximal's FrontierSWE) do not cover them. A published price answers a transparency question, not a trust one.
What changes for you
- Re-forecast every GLM-5.3-Flash workload at $0.15 input and $0.50 output per 1M tokens. If your number came from a launch write-up, double it.
- Budget full-size GLM-5.3 at $1.40 / $4.40 with $0.26 cached input. If you already priced GLM-5.2, nothing changes.
- Do not standardize on GLM-5.3's headline benchmark scores until someone outside Z.ai reproduces them.
FAQ
Is GLM-5.3-Flash still $0.075 per 1M input tokens? No. That was a 50 percent launch promotion that ended at 24:00 UTC+8 on 9 September 2026. The list rate is $0.15 input and $0.50 output per 1M tokens.
Did full-size GLM-5.3 get more expensive? No. It is $1.40 input and $4.40 output per 1M tokens, the same as GLM-5.2. Only the Flash tier moved.
Why is the verdict still conditional? Because Z.ai's own team ran the headline coding and cyber benchmarks, and no pricing fact changes that.
Rating unchanged: the headline coding and cyber benchmarks are still run by Z.ai's own team. The standing summary is corrected: it said per-token pricing was unpublished when Z.ai's table lists $1.40/$4.40, and the entity had stored the expired GLM-5.3-Flash launch promotion as if it were the list rate.
What to do
- 1 Re-forecast every GLM-5.3-Flash workload at $0.15 input and $0.50 output per 1M tokens, not the $0.075/$0.25 launch promotion that ended 9 September 2026.
- 2 Budget full-size GLM-5.3 at $1.40 / $4.40 per 1M tokens with $0.26 cached input; if you already priced GLM-5.2, nothing changes.
- 3 Do not standardize on GLM-5.3's headline benchmark scores until a third party reproduces them.
Affected tools & models
Never need to catch up again
The weekly delta — only verdict changes and act-now items. No digest filler.