nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16Nemotron 3.5 Lightning 30B is an LLM listed in RunInfra Model APIs.
API access is paused until . Inference requests will not run while the model is paused.
USD, pay per token
Hosted inference is available on every plan, including Free. You pay for token usage from your workspace balance.
Output speed
Time to first token
Decode at 10,000 token input
512.9 output tok/s
Measured August 13, 2026 on the serving host as a single stream with a 10,000 token input, 1,500 output tokens, and temperature 0, in BF16 on 2x NVIDIA H100 SXM5 with speculative decoding using 3 draft tokens.
This figure comes from a measured single-stream request run directly on the serving host for this endpoint. Measured August 13, 2026 on the serving host as a single stream with a 1,000 token input, 1,500 output tokens, and temperature 0, in BF16 on 2x NVIDIA H100 SXM5 with speculative decoding using 3 draft tokens.
Confirm how your client reaches this model.
Check the limits your workload must fit.
See which request modes the API supports.
Save these examples for when access returns. Set RUNINFRA_GATEWAY_KEY to your workspace API key before using them.
Verify the company and operating credentials behind this API.
© 2026 RunInfra. All rights reserved.