vLLM
The model card states a vLLM v0.23.0+ serving floor for GLM-5.2.
This reference covers the open flagship checkpoint and preserves conflicting source claims side by side. It separates vendor deployment floors from independent serving support.
GLM-5.2 is the vendor reference for zai-org/GLM-5.2, as of 2026-08-12. Vendor-stated scale: 744B parameters (40B active); Context length: Solid 1M-token context that stably sustains long-horizon work, as of 2026-08-12. RunInfra has not measured this model; every figure below belongs to its named source, cited and dated.
| Group | Fact | Source-cited display value | Source and date |
|---|---|---|---|
| Identity | Release identity | zai-org/GLM-5.2 | SourceRetrieved 2026-08-12 |
| Identity | API model id | glm-5.2 | SourceRetrieved 2026-08-12 |
| Identity | Family position | Newest GLM flagship listed by the organization at retrieval. | SourceRetrieved 2026-08-12 |
| Identity | Hugging Face card date | June 17, 2026 | SourceRetrieved 2026-08-12 |
| Architecture | Vendor-stated scale | 744B parameters (40B active) | SourceRetrieved 2026-08-12 |
| Architecture | Repository counter | The Hugging Face automatic parameter counter shows about 753B, while the vendor states 744B. | SourceRetrieved 2026-08-12 |
| Architecture | Sparse attention | DSA sparse attention with IndexShare, which reuses the same indexer across every four sparse attention layers, reducing per-token FLOPs by 2.9x at a 1M context length | SourceRetrieved 2026-08-12 |
| Architecture | Speculative decoding | MTP layer for speculative decoding, increasing the acceptance length by up to 20% | SourceRetrieved 2026-08-12 |
| Architecture | Configuration | 78 layers, 256 routed experts, 1 shared expert, 8 experts per token, max_position_embeddings 1048576, model_type glm_moe_dsa | SourceRetrieved 2026-08-12 |
| Context | Context length | Solid 1M-token context that stably sustains long-horizon work | SourceRetrieved 2026-08-12 |
| Context | Maximum output | 128K | SourceRetrieved 2026-08-12 |
| Context | Maximum output | maximum generation length of 163,840 tokens | SourceRetrieved 2026-08-12 |
| Modalities | Modality | text | SourceRetrieved 2026-08-12 |
| Modalities | Thinking control | Thinking toggle with reasoning_effort including max | SourceRetrieved 2026-08-12 |
| Modalities | Agentic support | Tool calling and MCP integration | SourceRetrieved 2026-08-12 |
| License | Weights license | MIT; Copyright (c) 2026 Zhipu AI | SourceRetrieved 2026-08-12 |
| Pricing | Official input | $1.40 per 1M tokens | SourceRetrieved 2026-08-12 |
| Pricing | Official cached input | $0.26 per 1M tokens | SourceRetrieved 2026-08-12 |
| Pricing | Official output | $4.40 per 1M tokens | SourceRetrieved 2026-08-12 |
| Availability | Repositories | zai-org/GLM-5.2 and zai-org/GLM-5.2-FP8 exist, alongside GLM-5.1 and GLM-5 siblings. | SourceRetrieved 2026-08-12 |
| Availability | Quantized repository downloads | zai-org/GLM-5.2-FP8 had 2.03M downloads at retrieval. | SourceRetrieved 2026-08-12 |
| Availability | Vendor serving floors | vLLM v0.23.0+ and SGLang v0.5.13.post1+ | SourceRetrieved 2026-08-12 |
| Availability | Reference weights | zai-org/GLM-5.2 | SourceRetrieved 2026-08-12 |
| Availability | Quantized weights | zai-org/GLM-5.2-FP8 | SourceRetrieved 2026-08-12 |
| Availability | Earlier sibling weights | zai-org/GLM-5.1 | SourceRetrieved 2026-08-12 |
| Availability | Earlier sibling weights | zai-org/GLM-5 | SourceRetrieved 2026-08-12 |
"IndexShare, which reuses the same indexer across every four sparse attention layers, reducing per-token FLOPs by 2.9x at a 1M context length"
SourceRetrieved 2026-08-12
"for speculative decoding, increasing the acceptance length by up to 20%"
SourceRetrieved 2026-08-12
"Stronger coding capabilities with multiple thinking effort levels"
SourceRetrieved 2026-08-12
These results are vendor-claimed, not independently measured by RunInfra.
RunInfra has not measured this model yet. When measurement is published, the record will state throughput, latency, memory, quality, serving conditions, and reproducible evidence.
As of 2026-08-12, release identity: zai-org/GLM-5.2.
As of 2026-08-12, repositories: zai-org/GLM-5.2 and zai-org/GLM-5.2-FP8 exist, alongside GLM-5.1 and GLM-5 siblings.
As of 2026-08-12, repositories: zai-org/GLM-5.2 and zai-org/GLM-5.2-FP8 exist, alongside GLM-5.1 and GLM-5 siblings.
© 2026 RunInfra. All rights reserved.