RunInfraby RightNow
  • Model APIsNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started

Model Library

The best hosted models for agentic workloads. Sessions stay on the replica that holds their prefix, so every replay bills at the cached input rate where published.

RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

Join the communitySystem status
Optimization agentModel APIsPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

Join the communitySystem status
Optimization agentModel APIsPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy
  1. Model Library
  2. /
  3. Z.ai
  4. /
  5. GLM 5.3 Flash

GLM 5.3 Flash

zai-org/GLM-5.3-Flash
Get API keyView docs

GLM 5.3 Flash is an LLM listed in RunInfra Model APIs. RunInfra serves it as zai-org/GLM-5.3-Flash at $0.10 per 1M input tokens and $0.40 per 1M output tokens. Its context window is 1,048,576 tokens. The API provides OpenAI-compatible chat completions.

Pricing

USD, pay per token

per 1M input tokens
$0.10
per 1M cached input tokens
$0.01
per 1M output tokens
$0.40

Measured performance

Output speed

254.1output tokens per second, model only

Time to first token

703milliseconds to first reasoning token

Access

Confirm how your client reaches this model.

Provider
Z.ai
API compatibility
OpenAI-compatible chat completions
Accepted input
Text only
Availability
Available
Data retention
Your prompts are never stored and never used for training.

Code examples

Set RUNINFRA_GATEWAY_KEY to your workspace API key before using an example.

Trust and provenance

Verify the company and operating credentials behind this API.

RightNowRunInfra is a sub-product of RightNow Research Lab.
  • SOC 2 Type IIAudited access, logging, and incident response.
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

Join the communitySystem status

Capacity

Check the limits your workload must fit.

Context window
1,048,576 tokens
Maximum request size
3.5 MB per request
Maximum generated output
1,048,576 tokens

Capabilities

See which request modes the API supports.

Tool calling
Supported
JSON mode
Supported
Streaming
Supported
Precision
FP8, vendor-native release
View Full Spec
Gateway compatibility
OpenAI-compatible chat completions for compatible clients and gateways
OpenRouter Provider Monitor
Provider metadata published as ready for discovery. Not proof of a live OpenRouter listing or callability.
Prefix caching
Automatic prefix caching runs on every replica, and requests from one workspace or session are routed back to the replica that holds their cached prefix. A stable session hint (the x-session-affinity header, prompt_cache_key, or user) strengthens the grouping. Hits remain best effort until eviction or a serve restart, never guaranteed. Cached input is billed at the cached input rate. Responses report cached token counts.
Cache retention
Prefix cache is held on the serving GPU and evicted under memory pressure. Retention is best effort.
Upstream model
View model
Y CombinatorBacked by Y Combinator.
  • NVIDIA InceptionMember of NVIDIA Inception.
  • Optimization agent
    Model APIs
    Pricing
    Startups
    Benchmarks
    Docs
    Research
    News
    Contact
    Backed by
    YCombinator
    AICPA Type II
    SOC 2
    NVIDIA Inception ProgramNVIDIA Inception Program
    Ask AI about RunInfra
    Part of RightNow
    SecurityDPAAUPCookiesTermsPrivacy