Claims and evidence
- Released Other sources give: [1]Published Publication date of the NVIDIA introduction article on Hugging Face. Primary source[2]Release Date
- Status AvailableNVIDIA has no deprecation page for its open models; the weights were still downloadable from the NVIDIA organisation on Hugging Face on 2026-10-01. Primary source[2]
- Derived from (distillation) Nemotron Nano 2Pruned from NVIDIA-Nemotron-Nano-9B-v2 with the Nemotron Elastic framework, then retrained by knowledge distillation from the frozen 9B parent. Primary source[1]Compressing 9B to 4B with Nemotron Elastic; Two-Stage Distillation for Accuracy Recovery [2]Model Overview; Model Architecture
- Change · Size Compressed from the 9B Nemotron Nano 2 model to about 4B parameters.Compared with Nemotron Nano 2 Primary source[1]Compressing 9B to 4B with Nemotron Elastic
- Change · Architecture Pruning cut depth from 56 to 42 layers and also reduced Mamba heads, feed-forward width and embedding size.Compared with Nemotron Nano 2 Primary source[1]Compressing 9B to 4B with Nemotron Elastic
- Change · Availability Built for on-device use on Jetson, GeForce RTX and DGX Spark; NVIDIA calls it its first model optimised for that purpose.Compared with Nemotron Nano 2 Primary source[1]
- Change · Training data Post-trained with a recipe derived from the Nemotron 3 post-training data, with reinforcement learning focused on instruction following and tool calling.Compared with Nemotron Nano 2 Primary source[1]
- Input text Primary source[2]Input
- Output text Primary source[2]Output
- Feature Reasoning modeOne model for reasoning and non-reasoning tasks; the reasoning trace can be turned off through the system prompt. Primary source[2]Model Overview
- Feature Tool use Primary source[1]
- Open weights YesReleased under the NVIDIA Nemotron Open Model License in BF16, FP8 and Q4_K_M GGUF. Primary source[2]License/Terms of Use [1]Try It Now!
- Context window 262K tokensStated as context length up to 262K. Primary source[2]Input
- Parameters 3.97 x 10^9The introduction article rounds this to 4 billion parameters. Primary source[2]Model Architecture
- Access Open-weights download, On devicePositioned for local deployment on Jetson Thor, Jetson Orin Nano, DGX Spark and RTX GPUs. Primary source[1]
Lineage
Predecessors
No known predecessor.
Successors
No known successor.
Based on
- Nemotron Nano 2 · 18 August 2025 · derived (distillation)
Variants and derived
None recorded.
Siblings
None recorded.
All ancestors
- LLaMA · Meta · 24 February 2023
- Llama 2 · Meta · 18 July 2023
- Meta Llama 3 · Meta · 18 April 2024
- Llama 3.1 · Meta · 23 July 2024
- Llama-3.1-Nemotron-Nano-8B-v1 · 18 March 2025
- Nemotron Nano 2 · 18 August 2025
All descendants
None.
Variants
No variants recorded in this record.
Related AI Radar coverage
AI Radar coverage starts in June 2026; no coverage linked yet.