Nemotron 3 Nano 4B

Available · Language, Reasoning

Nemotron 3 Nano 4B is a small open-weights text model in NVIDIA's Nemotron 3 family that handles reasoning and non-reasoning tasks. NVIDIA calls it its first model optimised for on-device use and built it by pruning and distilling Nemotron Nano 9B v2 into a hybrid of mostly Mamba-2 layers with about 4B parameters. [1] [2]Model Overview; Model Architecture Primary source

Timeline of Nemotron 3 Nano 4B →

Claims and evidence

  • Released Other sources give: [1]Published Publication date of the NVIDIA introduction article on Hugging Face. Primary source[2]Release Date
  • Status AvailableNVIDIA has no deprecation page for its open models; the weights were still downloadable from the NVIDIA organisation on Hugging Face on 2026-10-01. Primary source[2]
  • Derived from (distillation) Nemotron Nano 2Pruned from NVIDIA-Nemotron-Nano-9B-v2 with the Nemotron Elastic framework, then retrained by knowledge distillation from the frozen 9B parent. Primary source[1]Compressing 9B to 4B with Nemotron Elastic; Two-Stage Distillation for Accuracy Recovery [2]Model Overview; Model Architecture
  • Change · Size Compressed from the 9B Nemotron Nano 2 model to about 4B parameters.Compared with Nemotron Nano 2 Primary source[1]Compressing 9B to 4B with Nemotron Elastic
  • Change · Architecture Pruning cut depth from 56 to 42 layers and also reduced Mamba heads, feed-forward width and embedding size.Compared with Nemotron Nano 2 Primary source[1]Compressing 9B to 4B with Nemotron Elastic
  • Change · Availability Built for on-device use on Jetson, GeForce RTX and DGX Spark; NVIDIA calls it its first model optimised for that purpose.Compared with Nemotron Nano 2 Primary source[1]
  • Change · Training data Post-trained with a recipe derived from the Nemotron 3 post-training data, with reinforcement learning focused on instruction following and tool calling.Compared with Nemotron Nano 2 Primary source[1]
  • Input text Primary source[2]Input
  • Output text Primary source[2]Output
  • Feature Reasoning modeOne model for reasoning and non-reasoning tasks; the reasoning trace can be turned off through the system prompt. Primary source[2]Model Overview
  • Feature Tool use Primary source[1]
  • Open weights YesReleased under the NVIDIA Nemotron Open Model License in BF16, FP8 and Q4_K_M GGUF. Primary source[2]License/Terms of Use [1]Try It Now!
  • Context window 262K tokensStated as context length up to 262K. Primary source[2]Input
  • Parameters 3.97 x 10^9The introduction article rounds this to 4 billion parameters. Primary source[2]Model Architecture
  • Access Open-weights download, On devicePositioned for local deployment on Jetson Thor, Jetson Orin Nano, DGX Spark and RTX GPUs. Primary source[1]

Lineage

Predecessors

No known predecessor.

Successors

No known successor.

Based on

Variants and derived

None recorded.

Siblings

None recorded.

All ancestors

All descendants

None.

Variants

No variants recorded in this record.

Related AI Radar coverage

AI Radar coverage starts in June 2026; no coverage linked yet.

All model releases from NVIDIA on AI Radar →