Nemotron 3 Nano

Available · Language, Reasoning · Milestone

Nemotron 3 Nano is the first released model of NVIDIA's Nemotron 3 open model family: a hybrid Mamba-Transformer mixture-of-experts reasoning model with 31.6 billion total and 3.2 billion active parameters and a context of up to 1 million tokens. NVIDIA published the weights, training recipe and redistributable training data, and offered it through inference providers and as an NVIDIA NIM. [1] [2] Primary source

Timeline of Nemotron 3 Nano →

Claims and evidence

  • Announced Pre-announced at GTC DC as NVIDIA Nemotron Nano 3, a 32B-parameter MoE with 3.6B active parameters, available soon; the paragraph is in the Internet Archive capture of that day. The full Nemotron 3 family (Nano, Super, Ultra) was announced on 2025-12-15 together with the Nano release. Primary source[3]Enable agents to think efficiently with NVIDIA Nemotron Nano 3
  • Released Primary source[1]Get Started With NVIDIA Open Models [2]Published: December 15, 2025 [4]Release Date
  • Status AvailableNVIDIA has no deprecation page for its open models; the weights were still downloadable from NVIDIA's Hugging Face organisation on 2026-10-01. Primary source[4] [2]Open Source
  • Successor of Nemotron Nano 2NVIDIA calls Nemotron 2 Nano its previous generation and compares the two models. Primary source[2]Nemotron 3 Nano [5]Abstract
  • Change · Architecture Moves to a mixture-of-experts hybrid Mamba-Transformer design that activates less than half as many parameters per forward pass as Nemotron 2 Nano.Compared with Nemotron Nano 2 Primary source[2]
  • Change · Context length Context window grows from 128K tokens to up to 1M tokens.Compared with Nemotron Nano 2 Primary source[1] [6]Models
  • Change · Efficiency NVIDIA says token throughput is up to four times higher and reasoning-token generation up to 60 percent lower than with Nemotron 2 Nano.Compared with Nemotron Nano 2 Primary source[1]
  • Change · Training data Pretrained on 25 trillion text tokens, including more than 3 trillion new unique tokens compared with Nemotron 2.Compared with Nemotron Nano 2 Primary source[5]
  • Input text Primary source[4]Input
  • Output text Primary source[4]Output
  • Feature Reasoning modeReasoning on by default and switched off with enable_thinking=False in the chat template; a thinking budget can be set at inference time. Primary source[4]Description; Using Budget Control [7]Abstract
  • Feature Function callingThe card documents tool calling through a vLLM tool-call parser. Primary source[4]vLLM serving with tool-call parser
  • Feature Long context Primary source[2]Nemotron 3 technologies: Long Context
  • Open weights YesReleased on Hugging Face under the NVIDIA Nemotron Open Model License, together with the training recipe and the training data NVIDIA may redistribute. Primary source[2]Open Source [4]License/Terms of Use
  • Context window 1M tokensThe default context size in the Hugging Face configuration is 256k because of memory requirements. Primary source[4]Input; Output [1]Nemotron 3 Reinvents Multi-Agent AI [2]Nemotron 3 Nano
  • Parameters 31.6B total, 3.2B active (3.6B with embeddings)NVIDIA states this differently elsewhere: the press release says 30 billion parameters with up to 3 billion active, the model card 30B total and 3.5B active, and the October 2025 pre-announcement 32B with 3.6B active. Primary source[2]Nemotron 3 Nano
  • Access Open-weights download, API, cloud partnerAt release on Hugging Face, through inference providers such as Baseten, DeepInfra, Fireworks, FriendliAI, OpenRouter and Together AI, and as an NVIDIA NIM; Amazon Bedrock and other clouds were announced as coming soon. Primary source[1]Get Started With NVIDIA Open Models

Lineage

Predecessors

Successors

Based on

Not derived from another model.

Variants and derived

Siblings

None recorded.

All ancestors

All descendants

Variants

NVIDIA-Nemotron-3-Nano-30B-A3B

Same dates as Nemotron 3 Nano.

  • Variant Inline variant in this record.The post-trained model, released in BF16 and FP8; an NVFP4 checkpoint followed later (repository created 2025-12-20, Hugging Face API). Primary source[2]Open Source: Checkpoints [4]Release Date

Also known as: NVIDIA-Nemotron-3-Nano-30B-A3B-BF16, NVIDIA-Nemotron-3-Nano-30B-A3B-FP8, NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4, Nemotron 3 Nano 30B-A3B, Nemotron-3-Nano-30B-A3B

NVIDIA-Nemotron-3-Nano-30B-A3B-Base-BF16

Same dates as Nemotron 3 Nano.

  • Variant Inline variant in this record.The pre-trained base model before post-training. Primary source[2]Open Source: Checkpoints [8]Release Date

Also known as: Nemotron 3 Nano 30B-A3B Base

Related AI Radar coverage

AI Radar coverage starts in June 2026; no coverage linked yet.

All model releases from NVIDIA on AI Radar →