RunInfraby RightNow
  • Model APIsNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started

Model Library

Compare hosted model availability, prices, limits, and supported capabilities.

Loading the Model Library
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModel APIsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy
Loading Model APIs model
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModel APIsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy
  1. Model Library
  2. /
  3. NVIDIA
  4. /
  5. Nemotron 3.5 Lightning 30B

Nemotron 3.5 Lightning 30B

Paused
nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16
View API docs

Nemotron 3.5 Lightning 30B is an LLM listed in RunInfra Model APIs.

API access is paused until Sep 12, 2026, 11:07 AM UTC. Inference requests will not run while the model is paused.

Pricing

USD, pay per token

per 1M input tokens
$0.05
per 1M output tokens
$0.10

Hosted inference is available on every plan, including Free. You pay for token usage from your workspace balance.

Measured performance

Output speed

540.9output tokens per second

Time to first token

67milliseconds to first token

Decode at 10,000 token input

512.9 output tok/s

Measured August 13, 2026 on the serving host as a single stream with a 10,000 token input, 1,500 output tokens, and temperature 0, in BF16 on 2x NVIDIA H100 SXM5 with speculative decoding using 3 draft tokens.

This figure comes from a measured single-stream request run directly on the serving host for this endpoint. Measured August 13, 2026 on the serving host as a single stream with a 1,000 token input, 1,500 output tokens, and temperature 0, in BF16 on 2x NVIDIA H100 SXM5 with speculative decoding using 3 draft tokens.

Access

Confirm how your client reaches this model.

Provider
NVIDIA
API compatibility
OpenAI-compatible chat completions

Capacity

Check the limits your workload must fit.

Context window
262,144 tokens
Maximum request size
3.5 MB per request
Maximum generated output
32,768 tokens
Streaming time to first token limit
60 seconds
Maximum request duration
300 seconds

Capabilities

See which request modes the API supports.

Tool calling
Supported
JSON mode
Supported
Streaming
Supported
Served precision
BF16, unquantized
View Full Spec
Gateway compatibility
OpenAI-compatible chat completions for compatible clients and gateways
OpenRouter Provider Monitor
Provider metadata published, marked not ready while access is paused. Not proof of a live OpenRouter listing or callability.
Upstream model
View model

Code examples

Save these examples for when access returns. Set RUNINFRA_GATEWAY_KEY to your workspace API key before using them.

Trust and provenance

Verify the company and operating credentials behind this API.

RightNowRunInfra is a sub-product of RightNow Research Lab.
  • SOC 2 Type IIAudited access, logging, and incident response.
  • Y CombinatorBacked by Y Combinator.
  • NVIDIA InceptionMember of NVIDIA Inception.
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModel APIsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy