G

GLM-5.3

Z.ai · Released Aug 2026

Conditional

GLM-5.3 is Z.ai's latest open-weights coding and cyber-defense model — the same base as GLM-5.2, every gain from post-training. But the headline numbers are still vendor-run, the open weights are held ~2 weeks over emergent security capability, and per-token pricing is unpublished. Reach for defensive security and agentic automation now; wait for the weights and independent audits before standardizing.

Is it right for you?

Good for

  • Defensive cybersecurity — CyberGym 84.5%, top published result ahead of Mythos 5 (83.8%) and GPT-5.6 Sol (83.6%)
  • Long-horizon coding — DeepSWE v1.1 66.9% (from 46.2%), Terminal-Bench 3.0 28.3 (six-fold over 4.6)
  • Agentic automation — AutomationBench v1.0.6 jumps 26.2→48.2 (+84%); SAO reinforcement learning drives the long-horizon gains
  • Self-hosters planning open-weight adoption — reuses the GLM-5.2 base, weights staged ~2 weeks

Not good for

  • Deep offensive exploitation — ExploitBench 54.4% still trails Mythos 5 (78.0%) and GPT-5.6 Sol (76.5%)
  • Immediate self-hosting — weights held ~2 weeks for safety hardening; Z.ai's first cyber-motivated weight delay
  • Teams needing published per-token pricing or vision — no GLM-5.3 row on Z.ai's pricing table; no multimodal capability announced

How it performs by task

Defensive security (CyberGym)

Excellent

84.5% — best published result, ahead of the closed frontier; vendor-reported, not independently audited

Long-horizon software engineering (DeepSWE v1.1)

Very Good

66.9%, up from 46.2% — a 20-point jump, though the figure is Z.ai-run

Terminal/CLI coding (Terminal-Bench 3.0)

Very Good

28.3, six-fold over GLM-5.2's 4.6; level with the frontier on Terminal-Bench 2.1 (88.2 vs 88.8)

Deep exploitation (ExploitBench)

Fair

54.4% more than doubles GLM-5.2 (24.4%) but still trails the closed frontier by 20+ points

Vision / multimodal

Poor

No vision capability announced or benchmarked at launch

Pricing

Input

N/A (rate not yet published)

Output

N/A (rate not yet published)

Context

1M context

View full pricing

Benchmarks

BenchmarkScoreSource
CyberGym84.5% Source
DeepSWE v1.166.9% Source
Terminal-Bench 3.028.3 Source
Terminal-Bench 2.188.2 Source
ExploitBench54.4% Source
ExploitGym (2h / 6h tasks)105 / 130 Source

No verdict changes yet

The clock starts day one — changes land here as our verdict evolves.

Verification log

No verification checks yet

We haven't logged a verification check for this entry. Once a check runs, its history shows here.

How we evaluate