RunInfraby RightNow
  • Model APIsNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Loading related term
Loading related term
Loading related term
Loading related term
Home/Catalog/GLM-5.2
Model reference

GLM-5.2

This reference covers the open flagship checkpoint and preserves conflicting source claims side by side. It separates vendor deployment floors from independent serving support.

GLM-5.2 is the vendor reference for zai-org/GLM-5.2, as of 2026-08-12. Vendor-stated scale: 744B parameters (40B active); Context length: Solid 1M-token context that stably sustains long-horizon work, as of 2026-08-12. RunInfra has not measured this model; every figure below belongs to its named source, cited and dated.

Cited specifications

GroupFactSource-cited display valueSource and date
IdentityRelease identityzai-org/GLM-5.2
SourceRetrieved 2026-08-12
IdentityAPI model idglm-5.2
SourceRetrieved 2026-08-12
IdentityFamily positionNewest GLM flagship listed by the organization at retrieval.
SourceRetrieved 2026-08-12
IdentityHugging Face card dateJune 17, 2026
SourceRetrieved 2026-08-12
ArchitectureVendor-stated scale744B parameters (40B active)
SourceRetrieved 2026-08-12
ArchitectureRepository counterThe Hugging Face automatic parameter counter shows about 753B, while the vendor states 744B.
SourceRetrieved 2026-08-12
ArchitectureSparse attentionDSA sparse attention with IndexShare, which reuses the same indexer across every four sparse attention layers, reducing per-token FLOPs by 2.9x at a 1M context length
SourceRetrieved 2026-08-12
ArchitectureSpeculative decodingMTP layer for speculative decoding, increasing the acceptance length by up to 20%
SourceRetrieved 2026-08-12
ArchitectureConfiguration78 layers, 256 routed experts, 1 shared expert, 8 experts per token, max_position_embeddings 1048576, model_type glm_moe_dsa
SourceRetrieved 2026-08-12
ContextContext lengthSolid 1M-token context that stably sustains long-horizon work
SourceRetrieved 2026-08-12
ContextMaximum output128K
SourceRetrieved 2026-08-12
ContextMaximum outputmaximum generation length of 163,840 tokens
SourceRetrieved 2026-08-12
ModalitiesModalitytext
SourceRetrieved 2026-08-12
ModalitiesThinking controlThinking toggle with reasoning_effort including max
SourceRetrieved 2026-08-12
ModalitiesAgentic supportTool calling and MCP integration
SourceRetrieved 2026-08-12
LicenseWeights licenseMIT; Copyright (c) 2026 Zhipu AI
SourceRetrieved 2026-08-12
PricingOfficial input$1.40 per 1M tokens
SourceRetrieved 2026-08-12
PricingOfficial cached input$0.26 per 1M tokens
SourceRetrieved 2026-08-12
PricingOfficial output$4.40 per 1M tokens
SourceRetrieved 2026-08-12
AvailabilityRepositorieszai-org/GLM-5.2 and zai-org/GLM-5.2-FP8 exist, alongside GLM-5.1 and GLM-5 siblings.
SourceRetrieved 2026-08-12
AvailabilityQuantized repository downloadszai-org/GLM-5.2-FP8 had 2.03M downloads at retrieval.
SourceRetrieved 2026-08-12
AvailabilityVendor serving floorsvLLM v0.23.0+ and SGLang v0.5.13.post1+
SourceRetrieved 2026-08-12
AvailabilityReference weightszai-org/GLM-5.2
SourceRetrieved 2026-08-12
AvailabilityQuantized weightszai-org/GLM-5.2-FP8
SourceRetrieved 2026-08-12
AvailabilityEarlier sibling weightszai-org/GLM-5.1
SourceRetrieved 2026-08-12
AvailabilityEarlier sibling weightszai-org/GLM-5
SourceRetrieved 2026-08-12

What the vendor says is new

"IndexShare, which reuses the same indexer across every four sparse attention layers, reducing per-token FLOPs by 2.9x at a 1M context length"

SourceRetrieved 2026-08-12

"for speculative decoding, increasing the acceptance length by up to 20%"

SourceRetrieved 2026-08-12

"Stronger coding capabilities with multiple thinking effort levels"

SourceRetrieved 2026-08-12

Vendor-claimed benchmarks

These results are vendor-claimed, not independently measured by RunInfra.

BenchmarkVendor-claimed valueSource and date
Terminal-Bench 2.181.0
SourceRetrieved 2026-08-12
SWE-bench Pro62.1
SourceRetrieved 2026-08-12
GPQA-Diamond91.2
SourceRetrieved 2026-08-12
AIME 202699.2
SourceRetrieved 2026-08-12
HLE40.5
SourceRetrieved 2026-08-12

Serving support

Listed rows have a cited upstream support signal. They are not RunInfra measurements.

vLLM

The model card states a vLLM v0.23.0+ serving floor for GLM-5.2.

SourceRetrieved 2026-08-12

SGLang

The SGLang project publishes a GLM-5.2 serving cookbook.

SourceRetrieved 2026-08-12

Serving concepts

  • Mixture-of-experts serving->
  • Speculative decoding->
  • Expert parallelism->
  • Context length->
  • KV cache->

What RunInfra measures when we measure it

RunInfra has not measured this model yet. When measurement is published, the record will state throughput, latency, memory, quality, serving conditions, and reproducible evidence.

  • Measurement methodology->
  • Measured package catalog->
  • Published benchmarks->

Questions about this model reference

What is GLM-5.2?

As of 2026-08-12, release identity: zai-org/GLM-5.2.

Where are the GLM-5.2 weights?

As of 2026-08-12, repositories: zai-org/GLM-5.2 and zai-org/GLM-5.2-FP8 exist, alongside GLM-5.1 and GLM-5 siblings.

Can I run GLM-5.2 myself?

As of 2026-08-12, repositories: zai-org/GLM-5.2 and zai-org/GLM-5.2-FP8 exist, alongside GLM-5.1 and GLM-5 siblings.

RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModel APIsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy