GLM-5.3: Z.ai's Model Tops CyberGym, Weights in Two Weeks

Z.ai logoZ.aiImportantAugust 14, 2026Models
What happened
Z.ai released GLM-5.3, post-trained on the same base as GLM-5.2, with a more-than-6× Terminal-Bench 3.0 jump and CyberGym SOTA.
Why it matters
It leads CyberGym at 84.5%, ahead of Mythos 5 and GPT-5.6 Sol — a sharper open-weight frontier on coding and security.
What to do
Test via GLM Coding Plan or ZCode now; open weights land in about two weeks after safety review.

GLM-5.3 is the first open-weights model to lead CyberGym — and it beats Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol to get there. That's the headline. The catch: the weights don't ship today — they're staged about two weeks out, gated on safety evaluation and hardening. (Every number here is Z.ai's own, run through the Claude Code harness; we haven't independently reproduced them.)

What happened

Z.ai released GLM-5.3 as a post-training-only update on the same base model as GLM-5.2. "Scaling post-training is all we did," per Z.ai. This isn't a new model — it's what a month of environment-scaled RL bought.

BenchmarkGLM-5.2GLM-5.3Where it lands
Terminal-Bench 3.04.6%28.3%Open-source SOTA
Agents' Last Exam23.8%28.5%
CyberGym84.5%Ahead of Mythos 5 (83.8%), GPT-5.6 Sol (83.6%)
ExploitBench24.4%54.4%Still trails Mythos 5 (78.0%), GPT-5.6 Sol (76.5%)

Z.ai says the model found 2,436 vulnerabilities across 269 projects — 1,097 critical or high, some dating back to 1981.

Why it matters

The cyber capability is the part to watch. GLM-5.3 is the first open-weights model to top CyberGym, and the more-than-doubling on ExploitBench (24.4% to 54.4%) is the biggest gain exactly where the closed frontier is furthest ahead — it still trails Mythos 5 and GPT-5.6 Sol, but it closes a lot of that gap in a single post-training run.

For the open-weight world, this sharpens the frontier on both coding and security. But "open" is a promise with a two-week fuse, not a fact today: right now the only way to reach GLM-5.3 is through the GLM Coding Plan (usable in ZCode, Claude Code, OpenCode).

What changes for you

  • You can't run it locally yet. Weights land on Hugging Face in about two weeks, after the safety review and hardening.
  • A breaking change ships today. thinking.type: "disabled" is no longer supported — set thinking.type to "enabled" and use reasoning_effort (low/high/max).
  • Test it now through the GLM Coding Plan or ZCode if you want the capability before the weights drop.

FAQ

Is GLM-5.3 actually open weights today? No. The weights are staged roughly two weeks out, gated on safety evaluation and hardening. Until then it's only reachable via the GLM Coding Plan and ZCode.

Is this a new model? No — it's post-training on the same base as GLM-5.2. The gains come from scaling post-training alone, which is what makes the Terminal-Bench 3.0 jump from 4.6% to 28.3% notable.

What do I change if I migrate from GLM-5.2? Set thinking.type to "enabled" and add reasoning_effort (low/high/max). thinking.type: "disabled" no longer works.

Bottom line: the open-weight frontier just got sharper on both coding and security — but the actual weights are a two-week promise, not something you can download today.

What to do

  1. 1 Watch for GLM-5.3 weights on Hugging Face in about two weeks (gated on safety evaluation and hardening).
  2. 2 Test now via GLM Coding Plan or ZCode — note thinking.type: "disabled" is gone — set it to "enabled" and use reasoning_effort.

Affected tools & models

GLM-5.2GLM-5.3

Never need to catch up again

The weekly delta — only verdict changes and act-now items. No digest filler.

By subscribing you agree to our Privacy Policy. Unsubscribe anytime.