Skip to content

GLM 5.3 Flash

zai-org/GLM-5.3-Flash

GLM 5.3 Flash is an LLM listed in RunInfra Model APIs. RunInfra serves it as zai-org/GLM-5.3-Flash at $0.08 per 1M input tokens until Oct 13, 2026, 5:44 AM UTC, with standard rates resuming automatically. Its context window is 1,048,576 tokens. The API provides OpenAI-compatible chat completions and Anthropic-compatible Messages, POST /v1/messages.

Pricing

USD, pay per token

Normally $0.11
per 1M input tokens
$0.08
Normally $0.03
per 1M cached input tokens
$0.02
Normally $0.45
per 1M output tokens
$0.33

Input and output are 26% off until .

02days
08hrs
28min
46sec
Ends in 2 days 8 hours

Standard rates resume automatically when the window ends.

Measured performance

Output speed

254.1output tokens per second, model only

Time to first token

703milliseconds to first reasoning token

Cache hit rate

99.8%last 24 hours

Access

Confirm how your client reaches this model.

Provider
Z.ai
API compatibility
OpenAI-compatible chat completions
Anthropic compatibility
Anthropic-compatible Messages, POST /v1/messages
Accepted input
Text and images
Availability
Available

Capacity

Check the limits your workload must fit.

Context window
1,048,576 tokens
Maximum request size
3.5 MB per request
Maximum generated output
1,048,576 tokens

Capabilities

See which request modes the API supports.

Tool calling
Supported
JSON mode
Supported
Streaming
Supported
Precision
FP8, vendor-native release
View full spec
Gateway compatibility
OpenAI- and Anthropic-compatible APIs for compatible clients, tools, and gateways
Prefix caching
Automatic prefix caching runs on every replica, and requests from the same session or conversation are routed back to the replica that holds their cached prefix. A stable session hint or prompt_cache_key provides explicit grouping; otherwise RunInfra derives it from the start of the conversation or, failing that, the workspace. Hits remain best effort until eviction or a serve restart, never guaranteed. Cached input is billed at the cached input rate. Responses report cached token counts.
Cache retention
Prefix cache is held on the serving GPU and evicted under memory pressure. Retention is best effort.
Upstream model
View model