GLM 5.3 Flash
zai-org/GLM-5.3-FlashGLM 5.3 Flash is an LLM listed in RunInfra Model APIs. RunInfra serves it as zai-org/GLM-5.3-Flash at $0.08 per 1M input tokens until Oct 13, 2026, 5:44 AM UTC, with standard rates resuming automatically. Its context window is 1,048,576 tokens. The API provides OpenAI-compatible chat completions and Anthropic-compatible Messages, POST /v1/messages.
Pricing
USD, pay per token
- per 1M input tokens
- $0.08
- per 1M cached input tokens
- $0.02
- per 1M output tokens
- $0.33
Input and output are 26% off until .
Standard rates resume automatically when the window ends.
Measured performance
Output speed
Time to first token
Cache hit rate
Access
Confirm how your client reaches this model.
- Provider
- Z.ai
- API compatibility
- OpenAI-compatible chat completions
- Anthropic compatibility
- Anthropic-compatible Messages, POST /v1/messages
- Accepted input
- Text and images
- Availability
- Available
- Data retention
- Zero data retention by default. Never used for training.
Capacity
Check the limits your workload must fit.
- Context window
- 1,048,576 tokens
- Maximum request size
- 3.5 MB per request
- Maximum generated output
- 1,048,576 tokens
Capabilities
See which request modes the API supports.
- Tool calling
- Supported
- JSON mode
- Supported
- Streaming
- Supported
- Precision
- FP8, vendor-native release
View full spec
- Gateway compatibility
- OpenAI- and Anthropic-compatible APIs for compatible clients, tools, and gateways
- Prefix caching
- Automatic prefix caching runs on every replica, and requests from the same session or conversation are routed back to the replica that holds their cached prefix. A stable session hint or prompt_cache_key provides explicit grouping; otherwise RunInfra derives it from the start of the conversation or, failing that, the workspace. Hits remain best effort until eviction or a serve restart, never guaranteed. Cached input is billed at the cached input rate. Responses report cached token counts.
- Cache retention
- Prefix cache is held on the serving GPU and evicted under memory pressure. Retention is best effort.
- Upstream model
- View model