Nemotron 3 Nano Omni

Available · Language, Multimodal, Reasoning

Nemotron 3 Nano Omni is an open-weights multimodal reasoning model in NVIDIA's Nemotron 3 family that takes video, audio, images and text and produces text. It pairs the Nemotron 3 Nano 30B-A3B language model with vision and speech encoders; NVIDIA positions it as the perception sub-agent in agent systems and offered it as a download, on build.nvidia.com and via partners. [1] [2]Model Architecture: Network Architecture Primary source

Timeline of Nemotron 3 Nano Omni →

Claims and evidence

  • Released Primary source[1]At a Glance: Availability [2]Release Date [3]
  • Status AvailableNVIDIA has no deprecation page for its open models; the weights were still downloadable from the NVIDIA organisation on Hugging Face on 2026-10-01. Primary source[2]
  • Derived from (other) Nemotron 3 NanoUses the Nemotron 3 Nano 30B-A3B language model as its backbone, combined with the C-RADIOv4-H vision encoder and a Parakeet speech encoder. Primary source[2]Model Architecture: Network Architecture
  • Successor of Nemotron Nano 2 VLThe developer blog compares its multimodal accuracy with the previous Nemotron Nano VL V2 model. Primary source[3]Figure 3
  • Change · Modality Accepts audio input through a Parakeet speech encoder, alongside video, image and text input.Compared with Nemotron Nano 2 VL Primary source[2]Input(s)
  • Change · Architecture Uses the Nemotron 3 Nano 30B-A3B hybrid mixture-of-experts language model as its backbone, with the C-RADIOv4-H vision encoder.Compared with Nemotron Nano 2 VL Primary source[2]Model Architecture: Network Architecture
  • Change · Context length Context window of up to 256K tokens.Compared with Nemotron Nano 2 VL Primary source[2]At a Glance: Max context
  • Change · Licensing Released under the NVIDIA Open Model Agreement.Compared with Nemotron Nano 2 VL Primary source[2]License/Terms of Use
  • Input text, image, audio, video Primary source[2]Input(s) [1]At a Glance
  • Output text Primary source[2]Output(s)
  • Feature Reasoning modeOn by default; can be toggled. Primary source[2]At a Glance: Reasoning mode
  • Feature Tool useThe card says the model supports tool calling. Primary source[2]Output(s)
  • Open weights YesReleased under the NVIDIA Open Model Agreement in BF16, FP8 and NVFP4. Primary source[1]Open and Customizable, Deployable Anywhere [2]License/Terms of Use; Download Model Weights
  • Context window 256K tokensThe model card writes 256k tokens. Primary source[1]At a Glance: Architecture [2]At a Glance: Max context
  • Parameters 31B (~3B active per token)NVIDIA names the architecture 30B-A3B in the blogs; the model card gives 31B total parameters (3.1 x 10^10) with about 3B active per token. Primary source[2]At a Glance; Model Architecture
  • Access Open-weights download, API, cloud partnerHugging Face, OpenRouter, build.nvidia.com as an NVIDIA NIM microservice and partner platforms. Primary source[1]At a Glance: Availability; Open and Customizable, Deployable Anywhere

Lineage

Predecessors

Successors

No known successor.

Based on

Variants and derived

None recorded.

Siblings

None recorded.

All ancestors

All descendants

None.

Variants

No variants recorded in this record.

Related AI Radar coverage

AI Radar coverage starts in June 2026; no coverage linked yet.

All model releases from NVIDIA on AI Radar →