Nemotron Nano 2

Available · Language, Reasoning · Milestone

Nemotron Nano 2 is a family of open-weights hybrid Mamba-Transformer reasoning models that NVIDIA released on Hugging Face with 9B and 12B checkpoints and a 128K-token context. The 9B reasoning model, compressed from the 12B base so that 128K-token inference fits on one A10G GPU, was also on the NVIDIA API Catalog; users can switch reasoning off and cap thinking tokens. [1] [2] [3] Primary source

Timeline of Nemotron Nano 2 →

Claims and evidence

  • Released Primary source[1]Published: August 18, 2025; Models [2]Release Date
  • Status AvailableNVIDIA has no deprecation page for its open models; the weights of all checkpoints were still downloadable from NVIDIA's Hugging Face organisation on 2026-10-01. Primary source[2] [1]Models
  • Successor of Llama-3.1-Nemotron-Nano-8B-v1 EditorialEditorial link along the Nano tier: Nemotron Nano 2 is the next general-purpose Nano release and carries the number 2; NVIDIA pages do not name a predecessor. Unlike the Llama-based v1 model, it was trained from scratch by NVIDIA. Primary source[1]
  • Change · Architecture Hybrid Mamba-2 and Transformer design trained from scratch by NVIDIA, whereas the v1 Nano model was derived from a Llama model.Compared with Llama-3.1-Nemotron-Nano-8B-v1 Primary source[2]
  • Change · Reasoning Adds a thinking budget that can be set at inference time, next to switching reasoning on or off.Compared with Llama-3.1-Nemotron-Nano-8B-v1 Primary source[1]
  • Change · Training data Most of the pre-training data was released alongside the model as the Nemotron Pretraining Dataset v1 collection.Compared with Llama-3.1-Nemotron-Nano-8B-v1 Primary source[1]
  • Input text Primary source[2]Input
  • Output text Primary source[2]Output
  • Feature Reasoning modeReasoning switched on or off with /think and /no_think; the thinking budget can be set at inference time. Primary source[2]Model Overview; Reasoning Budget Control [1]Technical Highlights, Post-training
  • Feature Function calling Primary source[2]Using Tool-Calling with a vLLM Server
  • Open weights YesReleased on Hugging Face under the NVIDIA Open Model License Agreement. Primary source[1]Models [2]License/Terms of Use; Release Date
  • Context window 128K tokens Primary source[1]Models [2]Input
  • Parameters 8.89BCount for the selected pruned architecture of the 9B models (Candidate 2, which the report says was used for Nano 2); the official model name rounds it to 9B. The 12B variants have their own figure. Primary source[3]Section 4.2, Table 10 (Candidate 2)
  • Access Open-weights download, APIHugging Face and the NVIDIA API Catalog (build.nvidia.com) on 08/18/2025. Primary source[2]Release Date

Lineage

Predecessors

Successors

Based on

Not derived from another model.

Variants and derived

Siblings

None recorded.

All ancestors

All descendants

Variants

NVIDIA-Nemotron-Nano-9B-v2

Same dates as Nemotron Nano 2.

  • Variant Inline variant in this record.The aligned and pruned reasoning model, offered on Hugging Face and the NVIDIA API Catalog. Primary source[1]Models [2]Release Date

Also known as: Nemotron-Nano-9B-v2, NVIDIA-Nemotron-Nano-v2-9B

NVIDIA-Nemotron-Nano-9B-v2-Base

Same dates as Nemotron Nano 2.

  • Variant Inline variant in this record.Pruned base model, pruned and distilled from the 12B base model. Primary source[1]Models [4]Model Overview

Also known as: Nemotron-Nano-9B-v2-Base

NVIDIA-Nemotron-Nano-12B-v2-Base

Same dates as Nemotron Nano 2.

  • Variant Inline variant in this record.The base model before alignment or pruning, pre-trained on 20 trillion tokens in FP8. Primary source[1]Models [5]

Also known as: Nemotron-Nano-12B-v2-Base

Differs in:

  • Parameters: c. 12-billion-parameter Evidence not assessed [3]Abstract

NVIDIA-Nemotron-Nano-12B-v2

  • Released Primary source[6]Release Date
  • Variant Inline variant in this record.Aligned 12B reasoning checkpoint published eleven days after the first three; NVIDIA builds Nemotron Nano 2 VL on it. Primary source[6]

Also known as: Nemotron-Nano-12B-v2

Differs in:

  • Parameters: c. 12B Evidence not assessed [6]model name

Related AI Radar coverage

AI Radar coverage starts in June 2026; no coverage linked yet.

All model releases from NVIDIA on AI Radar →