RunInfraby RightNow
  • Model APIsNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Loading related term
Loading related term
Loading related term
Loading related term
Home/Catalog/NVIDIA Nemotron 3.5 Lightning 30B-A3B
Model reference

NVIDIA Nemotron 3.5 Lightning 30B-A3B

This reference covers the open Lightning checkpoint and its confirmed companion repositories. It separates vendor deployment guidance from measured RunInfra evidence.

NVIDIA Nemotron 3.5 Lightning 30B-A3B is the vendor reference for nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16, as of 2026-08-12. Scale: 30B Total / 3B Active; Context: Up to 1M tokens (for single H100 deployment, we use 256K), as of 2026-08-12. RunInfra has not measured this model; every figure below belongs to its named source, cited and dated.

Cited specifications

GroupFactSource-cited display valueSource and date
IdentityRelease identitynvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16
SourceRetrieved 2026-08-12
IdentityReleaseNemotron 3.5 Lightning; August 11, 2026
SourceRetrieved 2026-08-12
IdentityFamily positionan expansion of the Nemotron 3 model family, described by the vendor as the highest-efficiency model in its class for long-running agentic AI workloads
SourceRetrieved 2026-08-12
IdentityNIM API idnvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4
SourceRetrieved 2026-08-12
ArchitectureScale30B Total / 3B Active
SourceRetrieved 2026-08-12
ArchitectureArchitectureMixture-of-Experts Hybrid (Mamba + Transformer)
SourceRetrieved 2026-08-12
ArchitectureLayer designinterleaved Mamba-2 and MoE layers, along with select Attention layers
SourceRetrieved 2026-08-12
ArchitectureTraining designNemotron-3-Lightning + Multi-Token Prediction (MTP)
SourceRetrieved 2026-08-12
ArchitecturePre-trainingpre-trained with over 20T tokens
SourceRetrieved 2026-08-12
ContextContextUp to 1M tokens (for single H100 deployment, we use 256K)
SourceRetrieved 2026-08-12
ContextMaximum outputNot published by the vendor.
SourceRetrieved 2026-08-12
ModalitiesModalitytext
SourceRetrieved 2026-08-12
ModalitiesReasoningConfigurable on/off via chat template (enable_thinking=True/False)
SourceRetrieved 2026-08-12
ModalitiesStructured usetool calling, instruction following, structured outputs
SourceRetrieved 2026-08-12
ModalitiesLanguagesEnglish (and coding languages), Spanish, French, German, Italian, Japanese
SourceRetrieved 2026-08-12
LicenseLicenseOpenMDW License Agreement, version 1.1
SourceRetrieved 2026-08-12
LicenseLicense distinctionThis is not the NVIDIA Open Model License.
SourceRetrieved 2026-08-12
PricingNVIDIA pricingNot published by NVIDIA.
SourceRetrieved 2026-08-12
PricingOpenRouter paid input$0.05 per 1M input tokens
SourceRetrieved 2026-08-12
PricingOpenRouter paid output$0.20 per 1M output tokens
SourceRetrieved 2026-08-12
PricingOpenRouter free endpointFree endpoint listed.
SourceRetrieved 2026-08-12
AvailabilityReference weightsnvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16
SourceRetrieved 2026-08-12
AvailabilityCompanion repositoriesNVFP4, NVFP4-DSpark, and NVFP4-DFlash repositories are available.
SourceRetrieved 2026-08-12
AvailabilitySingle-device guidance1x H100 80GB (or 1x A100 80GB) for 256K context
SourceRetrieved 2026-08-12
AvailabilityFull-context guidance8x H100 - TP8 + expert parallel for 1M context
SourceRetrieved 2026-08-12
AvailabilityHardware familiesBlackwell (GB200, B200), Hopper (H100, H200), Ampere (A100)
SourceRetrieved 2026-08-12
AvailabilityReference weightsnvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16
SourceRetrieved 2026-08-12
AvailabilityBase weightsnvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-Base-BF16
SourceRetrieved 2026-08-12
AvailabilityProduction quantized weightsnvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4
SourceRetrieved 2026-08-12
AvailabilitySpeculative decoding draft variantnvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4-DSpark
SourceRetrieved 2026-08-12
AvailabilitySpeculative decoding draft variantnvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4-DFlash
SourceRetrieved 2026-08-12

What the vendor says is new

"This release follows Nemotron 3 Nano and reflects NVIDIA's commitment to continually improving open models for greater accuracy and speed."

SourceRetrieved 2026-08-12

Vendor-claimed benchmarks

These results are vendor-claimed, not independently measured by RunInfra.

BenchmarkVendor-claimed valueSource and date
MMLU Pro81.94
SourceRetrieved 2026-08-12
GPQA Diamond75.44
SourceRetrieved 2026-08-12
SWE-bench Verified51.56
SourceRetrieved 2026-08-12
HLE11.72
SourceRetrieved 2026-08-12

Serving support

Listed rows have a cited upstream support signal. They are not RunInfra measurements.

SGLang

Nemotron 3.5 Lightning cookbook support was merged on 2026-08-11.

SourceRetrieved 2026-08-12

SGLang

BF16 recipes were merged on 2026-08-12.

SourceRetrieved 2026-08-12

vLLM

The vendor pins a specific vLLM container version on the model card and a merged pull request fixed the model family's multi-token prediction path; no release note names 3.5 Lightning yet.

SourceRetrieved 2026-08-12

Serving concepts

  • Mixture-of-experts serving->
  • Speculative decoding->
  • Context length->
  • Expert parallelism->
  • Tokens per second->
  • KV cache->

What RunInfra measures when we measure it

RunInfra has not measured this model yet. When measurement is published, the record will state throughput, latency, memory, quality, serving conditions, and reproducible evidence.

  • Measurement methodology->
  • Measured package catalog->
  • Published benchmarks->

Questions about this model reference

What is NVIDIA Nemotron 3.5 Lightning 30B-A3B?

As of 2026-08-12, release identity: nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16.

Where are the NVIDIA Nemotron 3.5 Lightning 30B-A3B weights?

As of 2026-08-12, reference weights: nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16.

Can I run NVIDIA Nemotron 3.5 Lightning 30B-A3B myself?

As of 2026-08-12, single-device guidance: 1x H100 80GB (or 1x A100 80GB) for 256K context.

RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModel APIsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy