Fast responses. Lower-cost cached context. One API key for the OpenAI and Anthropic SDKs.
Nemotron 3.5 Lightning 30B, 1 of 4
Lower cached input price (USD / 1M tokens)
Session-aware routing keeps repeat calls close to cached context.
One session, one replica.
Repeat calls prefer the replica holding their context.
No hint needed.
Provide a session ID, or let automatic affinity use the conversation when available.
Reported, billed as cached.
Cache hits appear in usage and are billed at the cached rate. Hits stay best effort.
Per 1M input tokens, at each model's measured cache hit rate.
Price per 1M input tokens, log scale
cheaper, 99.6% cache hit
cheaper, 95.3% cache hit
cheaper, 96.7% cache hit
Streams report cached tokens when usage is requested. Your usage page shows your own billed figures.
4 models
Model capabilities: Tool calling / JSON mode / Streaming / Image input
Effective input $0.011 per 1M at that hit rate Effective input rate $0.011 per 1M tokens, from the measured cache share on billed agent traffic in the last 24 hours.
Model capabilities: Tool calling / JSON mode / Streaming
Effective input $0.012 per 1M at that hit rate Effective input rate $0.012 per 1M tokens, from the measured cache share on billed agent traffic in the last 24 hours.
Model capabilities: Tool calling / JSON mode / Streaming / Image input
Effective input $0.013 per 1M at that hit rate Effective input rate $0.013 per 1M tokens, from the measured cache share on billed agent traffic in the last 24 hours.
Model capabilities: Tool calling / JSON mode / Streaming / Image input
Works with the OpenAI and Anthropic SDKs.