SGLang
Nemotron 3.5 Lightning cookbook support was merged on 2026-08-11.
This reference covers the open Lightning checkpoint and its confirmed companion repositories. It separates vendor deployment guidance from measured RunInfra evidence.
NVIDIA Nemotron 3.5 Lightning 30B-A3B is the vendor reference for nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16, as of 2026-08-12. Scale: 30B Total / 3B Active; Context: Up to 1M tokens (for single H100 deployment, we use 256K), as of 2026-08-12. RunInfra serves this model on the Model APIs with measured numbers published at https://runinfra.ai/inference-api/nemotron-3-5-lightning-30b; no measured optimization package is published for it, and every figure below belongs to its named source, cited and dated.
| Group | Fact | Source-cited display value | Source and date |
|---|---|---|---|
| Identity | Release identity | nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 | SourceRetrieved 2026-08-12 |
| Identity | Release | Nemotron 3.5 Lightning; August 11, 2026 | SourceRetrieved 2026-08-12 |
| Identity | Family position | an expansion of the Nemotron 3 model family, described by the vendor as the highest-efficiency model in its class for long-running agentic AI workloads | SourceRetrieved 2026-08-12 |
| Identity | NIM API id | nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 | SourceRetrieved 2026-08-12 |
| Architecture | Scale | 30B Total / 3B Active | SourceRetrieved 2026-08-12 |
| Architecture | Architecture | Mixture-of-Experts Hybrid (Mamba + Transformer) | SourceRetrieved 2026-08-12 |
| Architecture | Layer design | interleaved Mamba-2 and MoE layers, along with select Attention layers | SourceRetrieved 2026-08-12 |
| Architecture | Training design | Nemotron-3-Lightning + Multi-Token Prediction (MTP) | SourceRetrieved 2026-08-12 |
| Architecture | Pre-training | pre-trained with over 20T tokens | SourceRetrieved 2026-08-12 |
| Context | Context | Up to 1M tokens (for single H100 deployment, we use 256K) | SourceRetrieved 2026-08-12 |
| Context | Maximum output | Not published by the vendor. | SourceRetrieved 2026-08-12 |
| Modalities | Modality | text | SourceRetrieved 2026-08-12 |
| Modalities | Reasoning | Configurable on/off via chat template (enable_thinking=True/False) | SourceRetrieved 2026-08-12 |
| Modalities | Structured use | tool calling, instruction following, structured outputs | SourceRetrieved 2026-08-12 |
| Modalities | Languages | English (and coding languages), Spanish, French, German, Italian, Japanese | SourceRetrieved 2026-08-12 |
| License | License | OpenMDW License Agreement, version 1.1 | SourceRetrieved 2026-08-12 |
| License | License distinction | This is not the NVIDIA Open Model License. | SourceRetrieved 2026-08-12 |
| Pricing | NVIDIA pricing | Not published by NVIDIA. | SourceRetrieved 2026-08-12 |
| Pricing | OpenRouter paid input | $0.05 per 1M input tokens | SourceRetrieved 2026-08-12 |
| Pricing | OpenRouter paid output | $0.20 per 1M output tokens | SourceRetrieved 2026-08-12 |
| Pricing | OpenRouter free endpoint | Free endpoint listed. | SourceRetrieved 2026-08-12 |
| Availability | Reference weights | nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 | SourceRetrieved 2026-08-12 |
| Availability | Companion repositories | NVFP4, NVFP4-DSpark, and NVFP4-DFlash repositories are available. | SourceRetrieved 2026-08-12 |
| Availability | Single-device guidance | 1x H100 80GB (or 1x A100 80GB) for 256K context | SourceRetrieved 2026-08-12 |
| Availability | Full-context guidance | 8x H100 - TP8 + expert parallel for 1M context | SourceRetrieved 2026-08-12 |
| Availability | Hardware families | Blackwell (GB200, B200), Hopper (H100, H200), Ampere (A100) | SourceRetrieved 2026-08-12 |
| Availability | Reference weights | nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 | SourceRetrieved 2026-08-12 |
| Availability | Base weights | nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-Base-BF16 | SourceRetrieved 2026-08-12 |
| Availability | Production quantized weights | nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 | SourceRetrieved 2026-08-12 |
| Availability | Speculative decoding draft variant | nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4-DSpark | SourceRetrieved 2026-08-12 |
| Availability | Speculative decoding draft variant | nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4-DFlash | SourceRetrieved 2026-08-12 |
"This release follows Nemotron 3 Nano and reflects NVIDIA's commitment to continually improving open models for greater accuracy and speed."
SourceRetrieved 2026-08-12
These results are vendor-claimed, not independently measured by RunInfra.
Listed rows have a cited upstream support signal. They are not RunInfra measurements.
Nemotron 3.5 Lightning cookbook support was merged on 2026-08-11.
BF16 recipes were merged on 2026-08-12.
The vendor pins a specific vLLM container version on the model card and a merged pull request fixed the model family's multi-token prediction path; no release note names 3.5 Lightning yet.
RunInfra serves this model on the Model APIs, where its measured serving numbers are published. No measured optimization package is published for it yet; when one is, the record will state throughput, latency, memory, quality, serving conditions, and reproducible evidence.
As of 2026-08-12, release identity: nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16.
As of 2026-08-12, reference weights: nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16.
As of 2026-08-12, single-device guidance: 1x H100 80GB (or 1x A100 80GB) for 256K context.
Use a workspace API key and pay for input, cached input, and output tokens.
View Model APIs