GLM-5.3: Z.ai's Model Tops CyberGym, Weights in Two Weeks
- O que aconteceu
- Z.ai released GLM-5.3, post-trained on the same base as GLM-5.2, with a more-than-6× Terminal-Bench 3.0 jump and CyberGym SOTA.
- Porque é importante
- It leads CyberGym at 84.5%, ahead of Mythos 5 and GPT-5.6 Sol — a sharper open-weight frontier on coding and security.
- O que fazer
- Test via GLM Coding Plan or ZCode now; open weights land in about two weeks after safety review.
GLM-5.3 is the first open-weights model to lead CyberGym — and it beats Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol to get there. That's the headline. The catch: the weights don't ship today — they're staged about two weeks out, gated on safety evaluation and hardening. (Every number here is Z.ai's own, run through the Claude Code harness; we haven't independently reproduced them.)
What happened
Z.ai released GLM-5.3 as a post-training-only update on the same base model as GLM-5.2. "Scaling post-training is all we did," per Z.ai. This isn't a new model — it's what a month of environment-scaled RL bought.
| Benchmark | GLM-5.2 | GLM-5.3 | Where it lands |
|---|---|---|---|
| Terminal-Bench 3.0 | 4.6% | 28.3% | Open-source SOTA |
| Agents' Last Exam | 23.8% | 28.5% | — |
| CyberGym | — | 84.5% | Ahead of Mythos 5 (83.8%), GPT-5.6 Sol (83.6%) |
| ExploitBench | 24.4% | 54.4% | Still trails Mythos 5 (78.0%), GPT-5.6 Sol (76.5%) |
Z.ai says the model found 2,436 vulnerabilities across 269 projects — 1,097 critical or high, some dating back to 1981.
Why it matters
The cyber capability is the part to watch. GLM-5.3 is the first open-weights model to top CyberGym, and the more-than-doubling on ExploitBench (24.4% to 54.4%) is the biggest gain exactly where the closed frontier is furthest ahead — it still trails Mythos 5 and GPT-5.6 Sol, but it closes a lot of that gap in a single post-training run.
For the open-weight world, this sharpens the frontier on both coding and security. But "open" is a promise with a two-week fuse, not a fact today: right now the only way to reach GLM-5.3 is through the GLM Coding Plan (usable in ZCode, Claude Code, OpenCode).
What changes for you
- You can't run it locally yet. Weights land on Hugging Face in about two weeks, after the safety review and hardening.
- A breaking change ships today.
thinking.type: "disabled"is no longer supported — setthinking.typeto"enabled"and usereasoning_effort(low/high/max). - Test it now through the GLM Coding Plan or ZCode if you want the capability before the weights drop.
FAQ
Is GLM-5.3 actually open weights today? No. The weights are staged roughly two weeks out, gated on safety evaluation and hardening. Until then it's only reachable via the GLM Coding Plan and ZCode.
Is this a new model? No — it's post-training on the same base as GLM-5.2. The gains come from scaling post-training alone, which is what makes the Terminal-Bench 3.0 jump from 4.6% to 28.3% notable.
What do I change if I migrate from GLM-5.2? Set thinking.type to "enabled" and add reasoning_effort (low/high/max). thinking.type: "disabled" no longer works.
Bottom line: the open-weight frontier just got sharper on both coding and security — but the actual weights are a two-week promise, not something you can download today.
O que fazer
- 1 Watch for GLM-5.3 weights on Hugging Face in about two weeks (gated on safety evaluation and hardening).
- 2 Test now via GLM Coding Plan or ZCode — note thinking.type: "disabled" is gone — set it to "enabled" and use reasoning_effort.
Ferramentas e modelos afetados
Nunca mais precisas de te pôr a par
O resumo semanal — apenas mudanças de veredicto e ações urgentes. Sem enchimento.