GLM-5.2
Z.ai · Released Jun 2026
Top open-weight coding model with 1M-token context and MIT license — beats GPT-5.5 on multiple benchmarks at a fraction of cost, but still trails Claude Opus 4.8 on the hardest long-horizon tasks.
Is it right for you?
Good for
- Strongest open-weight coding model at launch — 62.1 on SWE-bench Pro, 81.0 on Terminal-Bench 2.1
- 1M-token context window — 5× larger than GLM-5.1; can process entire codebases in one prompt
- MIT license with open weights — fully self-hostable, no revenue restrictions, commercial use allowed
- Massive improvement over GLM-5.1 — Terminal-Bench 2.1 jumped from 62.0 to 81.0; DeepSWE from 18.0 to 46.2
- IndexShare architecture cuts per-token FLOPs by 2.9× at 1M context — makes long-context coding economically viable
- Dual reasoning-effort modes (High and Max) — trade speed for depth
Not good for
- Heavy token appetite — some reports of high token consumption for equivalent tasks vs direct API calls
- Training code and full technical report not released at launch — only inference weights under MIT
- Still trails Claude Opus 4.8 on hardest tasks (SWE-bench Pro: 62.1 vs 69.2; SWE-Marathon: 13.0 vs 26.0)
How it performs by task
Code generation
Top open-weight model; beats GPT-5.5 on SWE-bench Pro (62.1) and Terminal-Bench 2.1 (81.0)
Long-horizon coding agents
1M context handles entire codebases; strong FrontierSWE (74.4) and DeepSWE (46.2) but trails Claude Opus 4.8
Competitive programming
AIME 2026: 99.2; IMOAnswerBench: 91.0 — class-leading open-weight math/coding reasoning
General reasoning
GPQA-D: 91.2; HLE: 40.5 — competitive but trails Claude Opus 4.8 (93.6/49.8)
Pricing
Input
$1.40 / 1M tokens
Output
$4.40 / 1M tokens
Context
1M tokens
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| SWE-bench Pro | 62.1 | Source |
| Terminal-Bench 2.1 (Terminus-2) | 81.0 | Source |
| Terminal-Bench 2.1 (Best Harness, Claude Code) | 82.7 | Source |
| FrontierSWE | 74.4 | Source |
| DeepSWE | 46.2 | Source |
| AIME 2026 | 99.2 | Source |
| GPQA-Diamond | 91.2 | Source |
| HLE (Humanity's Last Exam) | 40.5 | Source |
| HLE w/ Tools | 54.7 | Source |
| SWE-Marathon | 13.0 | Source |
| MCP-Atlas Public Set | 76.8 | Source |
| PostTrainBench | 34.3 | Source |
No verdict changes yet
The clock starts day one — changes land here as our verdict evolves.
Sources
- Developer TechJun 2026
- FelloAI — GLM-5.2 Complete GuideJun 2026
- Codersera — GLM-5.2 Complete GuideJun 2026
- Labellerr — GLM-5.2 Beats GPT-5.5Jun 2026
- VentureBeat — GLM-5.2 Benchmark CoverageJun 2026
- llm-stats.com — GLM-5.2 Pricing & BenchmarksJun 2026
- BuildFastWithAI — GLM-5.2 ReviewJun 2026
- kilo.ai — GLM-5.2 Model PageJun 2026
Verification log
- Pricing— No changes
Automated agent
- Pricing— No changes
Automated agent
- Pricing— Needs attention
Automated agent
Source unavailable: z.ai/pricing returned HTTP 404. Stored pricing ($1.40/$4.40) retained.
- Pricing— No changes
Automated agent
Pricing unchanged: $1.40/$4.40 per 1M tokens, 1M context. z.ai/pricing returned 404 but aipricing.guru confirms rates as of Jun 28.
- Pricing— No changes
Automated agent
GLM-5.2 unchanged: $1.40/$4.40, 1M context. z.ai/pricing 404 but FriendliAI & Novita confirm rates.
- Pricing— No changes
Automated agent
- Pricing— No changes
Imported at launch
- Profile— No changes
Imported at launch