Nemotron 3.5 Lightning

Available · Language, Reasoning

Nemotron 3.5 Lightning is an open-weights reasoning language model with 30B total and 3B active parameters that NVIDIA added to the Nemotron 3 family after Nemotron 3 Nano. NVIDIA aims it at high-volume agent work such as tool calls and subagent delegation, and released it with the NeMo Switchyard routing library, as a download and on build.nvidia.com. [1] [2] Primary source

Timeline of Nemotron 3.5 Lightning →

Claims and evidence

  • Released Primary source[1] [3]Model Summary: Release Date; Release Date [2]
  • Status AvailableNVIDIA has no deprecation page for its open models; the weights were still downloadable from the NVIDIA organisation on Hugging Face on 2026-10-01. Primary source[3]
  • Successor of Nemotron 3 NanoThe launch post says the release follows Nemotron 3 Nano and expands the Nemotron 3 family. Primary source[1]second paragraph
  • Change · Architecture Includes multi-token prediction layers from training onward and ships with DFlash and DSpark draft models for speculative decoding.Compared with Nemotron 3 Nano Primary source[2]
  • Change · Training data Pre-trained on more than 20 trillion tokens with an NVFP4 recipe; the card gives May 2026 as the post-training data cutoff.Compared with Nemotron 3 Nano Primary source[3]
  • Change · Licensing Released under the OpenMDW License Agreement version 1.1.Compared with Nemotron 3 Nano Primary source[3]Model Summary: License
  • Change · Other Developed with contributions from Nemotron Coalition members, who NVIDIA says supplied evaluation methods, inference software and datasets.Compared with Nemotron 3 Nano Primary source[1]
  • Input text Primary source[3]Input
  • Output text Primary source[3]Output
  • Feature Reasoning modeReasoning can be switched on or off through the chat template. Primary source[3]Model Summary: Reasoning Mode
  • Feature MultilingualEnglish, Spanish, French, German, Italian and Japanese. Primary source[3]Model Summary: Supported Languages
  • Open weights YesReleased under the OpenMDW License Agreement, version 1.1. Primary source[2]Customize Nemotron 3.5 Lightning out of the box [3]Model Summary: License
  • Context window 1M tokensStated as up to 1M tokens; NVIDIA uses 256K for a single H100 deployment. Primary source[3]Model Summary: Context Length
  • Parameters 30B (3B active) Primary source[3]Model Summary: Total Parameters [2]
  • Access Open-weights download, API, cloud partnerHugging Face, ModelScope, OpenRouter, build.nvidia.com as an NVIDIA NIM microservice and partner platforms. Primary source[1]last paragraph

Lineage

Predecessors

Successors

No known successor.

Based on

Not derived from another model.

Variants and derived

None recorded.

Siblings

None recorded.

All ancestors

All descendants

None.

Variants

Nemotron 3.5 Lightning 30B-A3B

Same dates as Nemotron 3.5 Lightning.

  • Variant Inline variant in this record.Post-trained model in BF16 and as an NVFP4 quantized checkpoint. The separate DFlash and DSpark repositories hold speculative decoding draft checkpoints, not standalone models, and are not listed as aliases. Primary source[3]Description [2]Quantization

Also known as: NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16, NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4

Nemotron 3.5 Lightning 30B-A3B Base

Same dates as Nemotron 3.5 Lightning.

  • Variant Inline variant in this record.Pre-trained base checkpoint without post-training. Primary source[4]Description; Release Date

Also known as: NVIDIA-Nemotron-3.5-Lightning-30B-A3B-Base-BF16

Related AI Radar coverage

All model releases from NVIDIA on AI Radar →