Claims and evidence
- Status AvailableNVIDIA has no deprecation page for its open models; the weights were still downloadable from the NVIDIA organisation on Hugging Face on 2026-10-01. Primary source[3]
- Successor of Nemotron 3 NanoThe launch post says the release follows Nemotron 3 Nano and expands the Nemotron 3 family. Primary source[1]second paragraph
- Change · Architecture Includes multi-token prediction layers from training onward and ships with DFlash and DSpark draft models for speculative decoding.Compared with Nemotron 3 Nano Primary source[2]
- Change · Training data Pre-trained on more than 20 trillion tokens with an NVFP4 recipe; the card gives May 2026 as the post-training data cutoff.Compared with Nemotron 3 Nano Primary source[3]
- Change · Licensing Released under the OpenMDW License Agreement version 1.1.Compared with Nemotron 3 Nano Primary source[3]Model Summary: License
- Change · Other Developed with contributions from Nemotron Coalition members, who NVIDIA says supplied evaluation methods, inference software and datasets.Compared with Nemotron 3 Nano Primary source[1]
- Input text Primary source[3]Input
- Output text Primary source[3]Output
- Feature Reasoning modeReasoning can be switched on or off through the chat template. Primary source[3]Model Summary: Reasoning Mode
- Feature MultilingualEnglish, Spanish, French, German, Italian and Japanese. Primary source[3]Model Summary: Supported Languages
- Open weights YesReleased under the OpenMDW License Agreement, version 1.1. Primary source[2]Customize Nemotron 3.5 Lightning out of the box [3]Model Summary: License
- Context window 1M tokensStated as up to 1M tokens; NVIDIA uses 256K for a single H100 deployment. Primary source[3]Model Summary: Context Length
- Parameters 30B (3B active) Primary source[3]Model Summary: Total Parameters [2]
- Access Open-weights download, API, cloud partnerHugging Face, ModelScope, OpenRouter, build.nvidia.com as an NVIDIA NIM microservice and partner platforms. Primary source[1]last paragraph
Lineage
Predecessors
- Nemotron 3 Nano · 15 December 2025
Successors
No known successor.
Based on
Not derived from another model.
Variants and derived
None recorded.
Siblings
None recorded.
All ancestors
- LLaMA · Meta · 24 February 2023
- Llama 2 · Meta · 18 July 2023
- Meta Llama 3 · Meta · 18 April 2024
- Llama 3.1 · Meta · 23 July 2024
- Llama-3.1-Nemotron-Nano-8B-v1 · 18 March 2025
- Nemotron Nano 2 · 18 August 2025
- Nemotron 3 Nano · 15 December 2025
All descendants
None.
Variants
Nemotron 3.5 Lightning 30B-A3B
Same dates as Nemotron 3.5 Lightning.
- Variant Inline variant in this record.Post-trained model in BF16 and as an NVFP4 quantized checkpoint. The separate DFlash and DSpark repositories hold speculative decoding draft checkpoints, not standalone models, and are not listed as aliases. Primary source[3]Description [2]Quantization
Also known as: NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16, NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4
Nemotron 3.5 Lightning 30B-A3B Base
Same dates as Nemotron 3.5 Lightning.
- Variant Inline variant in this record.Pre-trained base checkpoint without post-training. Primary source[4]Description; Release Date
Also known as: NVIDIA-Nemotron-3.5-Lightning-30B-A3B-Base-BF16
Related AI Radar coverage
- Dense vs. MoE Models: Active Parameters, Throughput, and When to Choose EachNVIDIA · Developer ·