Nemotron Nano 2 VL

Available · Language, Multimodal, Reasoning

Nemotron Nano 2 VL is a 12B open-weights vision-language model from NVIDIA that takes text, images and video and produces text, built on the Nemotron Nano 2 12B reasoning model with a RADIO vision encoder. NVIDIA presents it for document, multi-image and video understanding; it introduced Efficient Video Sampling and was offered on Hugging Face, build.nvidia.com and as an NVIDIA NIM. [1] [2] [3] Primary source

Timeline of Nemotron Nano 2 VL →

Claims and evidence

  • Released Primary source[2]Release Date [1]Add multimodal understanding and reasoning with NVIDIA Nemotron Nano 2 VL
  • Status AvailableNVIDIA has no deprecation page for its open models; the BF16, FP8 and NVFP4-QAD weights were still downloadable from NVIDIA's Hugging Face organisation on 2026-10-01. Primary source[2]
  • Derived from (other) Nemotron Nano 2Vision-language model built on the Nemotron Nano V2 12B reasoning LLM (NVIDIA-Nemotron-Nano-12B-v2) combined with a RADIO vision encoder and an MLP projector, trained in several supervised fine-tuning stages. Primary source[3]Abstract; 1 Introduction; 2 Model Architecture [2]Model Architecture
  • Change · Modality Adds image and video input to the text-only Nemotron Nano 2 language model it is built on.Compared with Nemotron Nano 2 Primary source[2] [4]Input
  • Change · Efficiency Introduces Efficient Video Sampling, which drops video patches that stay unchanged over time to cut token counts for long videos.Compared with Nemotron Nano 2 Primary source[3]
  • Input text, image, videoUp to four input images; video frames sampled at 2 FPS, 8 to 128 frames. Primary source[2]Input
  • Output text Primary source[2]Output
  • Feature Reasoning modeSupports reasoning-on and reasoning-off modes. Primary source[3]1 Introduction; 4.3 Reasoning Budget Control
  • Open weights YesCheckpoints in BF16, FP8 and NVFP4 (QAD) under the NVIDIA Open Model License Agreement. Primary source[2]License/Terms of Use; Release Date [3]Abstract
  • Context window 128K tokensStated as input plus output tokens. Primary source[2]Input: Input + Output Token [3]1 Introduction
  • Parameters 12.6BNVIDIA's blog and technical report round this to 12B. Primary source[2]Number of model parameters
  • Access Open-weights download, APIHugging Face and build.nvidia.com; the blog also says it is available as an NVIDIA NIM. Primary source[2]Release Date [1]Add multimodal understanding and reasoning with NVIDIA Nemotron Nano 2 VL

Lineage

Predecessors

No known predecessor.

Successors

Based on

Variants and derived

None recorded.

Siblings

None recorded.

All ancestors

All descendants

Variants

No variants recorded in this record.

Related AI Radar coverage

AI Radar coverage starts in June 2026; no coverage linked yet.

All model releases from NVIDIA on AI Radar →