NVIDIA models

NVIDIA is a technology company headquartered in Santa Clara, California, founded in 1993 to bring 3D graphics to gaming and multimedia. It describes itself as a pioneer of accelerated computing and designs chips, systems and software for AI computing. It also develops the Nemotron family of open models and publishes their weights and training data on Hugging Face.

Chronology

  1. 2026

    5 models
    1. Released
      Nemotron Lightning · Language, Reasoning Nemotron 3.5 Lightning

      Nemotron 3.5 Lightning is an open-weights reasoning language model with 30B total and 3B active parameters that NVIDIA added to the Nemotron 3 family after Nemotron 3 Nano. NVIDIA aims it at…

      • Includes multi-token prediction layers from training onward and ships with DFlash and DSpark draft models for speculative decoding.
      • Pre-trained on more than 20 trillion tokens with an NVFP4 recipe; the card gives May 2026 as the post-training data cutoff.
      Primary source Available · 1 news item
    2. Released
      Nemotron Ultra · Language, Reasoning · Milestone Nemotron 3 Ultra

      Nemotron 3 Ultra is an open-weights reasoning language model and the Ultra size of NVIDIA's Nemotron 3 family, with 550B total and 55B active parameters and a context length of up to 1M tokens…

      • Uses a hybrid Mamba-attention LatentMoE architecture with multi-token prediction layers, pre-trained in NVFP4.
      • 550B total parameters, of which 55B are active per token.
      Primary source Available · 1 news item
    3. Released
      Nemotron Nano · Language, Multimodal, Reasoning Nemotron 3 Nano Omni

      Nemotron 3 Nano Omni is an open-weights multimodal reasoning model in NVIDIA's Nemotron 3 family that takes video, audio, images and text and produces text. It pairs the Nemotron 3 Nano 30B-A3B…

      • Accepts audio input through a Parakeet speech encoder, alongside video, image and text input.
      • Uses the Nemotron 3 Nano 30B-A3B hybrid mixture-of-experts language model as its backbone, with the C-RADIOv4-H vision encoder.
      Primary source Available · no news linked
    4. Released
      Nemotron Nano · Language, Reasoning Nemotron 3 Nano 4B

      Nemotron 3 Nano 4B is a small open-weights text model in NVIDIA's Nemotron 3 family that handles reasoning and non-reasoning tasks. NVIDIA calls it its first model optimised for on-device use and…

      • Compressed from the 9B Nemotron Nano 2 model to about 4B parameters.
      • Pruning cut depth from 56 to 42 layers and also reduced Mamba heads, feed-forward width and embedding size.
      Primary source Available · no news linked
    5. Released
      Nemotron Super · Language, Reasoning · Milestone Nemotron 3 Super

      Nemotron 3 Super is an open-weights reasoning language model in NVIDIA's Nemotron 3 family, with 120B total and 12B active parameters, a hybrid Mamba-Transformer mixture-of-experts design and a…

      • Uses a hybrid Mamba-Transformer latent mixture-of-experts design with multi-token prediction, activating 12B of its 120B parameters per token.
      • Supports a context window of up to 1M tokens.
      Primary source Available · no news linked
  2. 2025

    9 models
    1. Released
      Nemotron Nano · Language, Reasoning · Milestone Nemotron 3 Nano

      Nemotron 3 Nano is the first released model of NVIDIA's Nemotron 3 open model family: a hybrid Mamba-Transformer mixture-of-experts reasoning model with 31.6 billion total and 3.2 billion active…

      • Moves to a mixture-of-experts hybrid Mamba-Transformer design that activates less than half as many parameters per forward pass as Nemotron 2 Nano.
      • Context window grows from 128K tokens to up to 1M tokens.
      Primary source Available · no news linked
    2. Released
      Nemotron Nano · Language, Multimodal, Reasoning Nemotron Nano 2 VL

      Nemotron Nano 2 VL is a 12B open-weights vision-language model from NVIDIA that takes text, images and video and produces text, built on the Nemotron Nano 2 12B reasoning model with a RADIO vision…

      • Adds image and video input to the text-only Nemotron Nano 2 language model it is built on.
      • Introduces Efficient Video Sampling, which drops video patches that stay unchanged over time to cut token counts for long videos.
      Primary source Available · no news linked
    3. Released
      Nemotron Nano · Language, Reasoning · Milestone Nemotron Nano 2

      Nemotron Nano 2 is a family of open-weights hybrid Mamba-Transformer reasoning models that NVIDIA released on Hugging Face with 9B and 12B checkpoints and a 128K-token context. The 9B reasoning…

      • Hybrid Mamba-2 and Transformer design trained from scratch by NVIDIA, whereas the v1 Nano model was derived from a Llama model.
      • Adds a thinking budget that can be set at inference time, next to switching reasoning on or off.
      Primary source Available · no news linked
    4. Released
      Nemotron Super · Language, Reasoning Llama-3.3-Nemotron-Super-49B-v1.5

      Llama-3.3-Nemotron-Super-49B-v1.5 is an open-weights NVIDIA reasoning language model of about 49 billion parameters, derived from Meta's Llama-3.3-70B-Instruct and released on Hugging Face and as a…

      • Further post-trained on additional reasoning data; NVIDIA says this improves math, science, coding, function calling, instruction following and chat.
      • NVIDIA trained tool calling with iterative DPO stages, and the model card provides a tool-call parser for serving the model with vLLM.
      Primary source Available · no news linked
    5. Released
      Nemotron Nano · Language, Reasoning Llama-3.1-Nemotron-Nano-4B-v1.1

      Llama-3.1-Nemotron-Nano-4B-v1.1 is a 4B open-weights reasoning language model in NVIDIA's Llama Nemotron Nano line, fine-tuned from a Minitron-compressed Llama 3.1 8B base and also offered as an…

      • About half the parameters of Nano 8B v1: 4 billion instead of 8 billion.
      • Built on Llama-3.1-Minitron-4B-Width-Base, which NVIDIA pruned and distilled from Llama 3.1 8B, instead of directly on Llama-3.1-8B-Instruct.
      Primary source Available · no news linked
    6. Released
      Nemotron · Language Nemotron-H

      Nemotron-H is an NVIDIA series of hybrid Mamba-Transformer language models that replace most self-attention layers with Mamba-2 layers to lower inference cost. The 8B, 47B and 56B base checkpoints…

      Primary source Available · no news linked
    7. Released
      Nemotron Ultra · Language, Reasoning · Milestone Llama-3.1-Nemotron-Ultra-253B-v1

      Llama-3.1-Nemotron-Ultra-253B-v1 is a 253B open-weights reasoning language model from NVIDIA, the Ultra size of the Llama Nemotron family, released in April 2025 and also offered as a hosted preview…

      Primary source Available · no news linked
    8. Released
      Nemotron Nano · Language, Reasoning · Milestone Llama-3.1-Nemotron-Nano-8B-v1

      Llama-3.1-Nemotron-Nano-8B-v1 is an 8B open-weights reasoning language model from NVIDIA, derived from Meta's Llama-3.1-8B-Instruct and released at GTC 2025 as the Nano size of the Llama Nemotron…

      Primary source Available · no news linked
    9. Released
      Nemotron Super · Language, Reasoning · Milestone Llama-3.3-Nemotron-Super-49B-v1

      Llama-3.3-Nemotron-Super-49B-v1 is an open-weights reasoning language model from NVIDIA, released at GTC 2025 as the Super size of the Llama Nemotron family and also offered as a hosted API. NVIDIA…

      Primary source Available · no news linked
  3. 2024

    6 models
    1. Day not given
      Nemotron · Language Llama-3.1-Nemotron-70B-Instruct

      Llama-3.1-Nemotron-70B-Instruct is a language model from NVIDIA, trained from Meta's Llama-3.1-70B-Instruct with reinforcement learning from human feedback and NVIDIA's own reward model to make…

      Primary source Available · no news linked
    2. Released
      Nemotron · Language Llama-3.1-Nemotron-51B-Instruct

      Llama-3.1-Nemotron-51B-Instruct is a 51-billion-parameter chat model from NVIDIA, released on 23 September 2024 and derived from Meta's Llama-3.1-70B through block-wise distillation and neural…

      Primary source Available · no news linked
    3. Released
      Nemotron · Language Nemotron-4 4B Instruct

      Nemotron-4 4B Instruct is a small language model from NVIDIA, announced on 20 August 2024 for NVIDIA ACE, its game-character technology, as an NVIDIA NIM for cloud and on-device use. NVIDIA says it…

      • About 4 billion parameters by name, distilled from the 15-billion-parameter Nemotron-4 15B.
      • Pruned, distilled and INT4-quantised by NVIDIA for on-device inference using about 2 GB of VRAM.
      Primary source Available · no news linked
    4. Released
      Mistral · Language Mistral NeMo

      Mistral NeMo is a 12B-parameter language model that Mistral AI built with NVIDIA, released with base and instruct weights under Apache 2.0 and offered on la Plateforme and as an NVIDIA NIM. It has a…

      • Has 12B parameters, up from 7B in Mistral 7B, while keeping a standard architecture that Mistral says lets it replace Mistral 7B as a drop-in.
      • Uses a new tokenizer, Tekken, based on Tiktoken, in place of the SentencePiece tokenizer of earlier Mistral models.
      Primary source Available · no news linked Co-developed; listed under Mistral AI
    5. Released
      Nemotron · Language Nemotron-4 340B

      Nemotron-4 340B is a family of language models from NVIDIA (base, instruct and reward), released on 14 June 2024 as open weights on Hugging Face and the NGC catalog under the NVIDIA Open Model…

      • 340 billion parameters, against 15 billion for Nemotron-4 15B, with a similar architecture.
      • Uses the same data blend as Nemotron-4 15B but trains on 9 trillion tokens: the first 8 trillion plus 1 trillion tokens of continued pre-training.
      Primary source Available · no news linked
    6. Announced, release date unknown
      Nemotron · Language · Milestone Nemotron-4 15B

      Nemotron-4 15B is a 15-billion-parameter multilingual language model from NVIDIA, described in a technical report posted on arXiv on 26 February 2024. It was trained on 8 trillion tokens of English…

      • 15 billion parameters, up from 8 billion for Nemotron-3 8B.
      • Pre-trained on 8 trillion tokens covering 43 programming languages, against 3.8 trillion tokens and 37 programming languages for the Nemotron-3 8B base model.
      Primary source Research only · no news linked
  4. 2023

    1 model
    1. Released
      Nemotron · Language · Milestone Nemotron-3 8B

      Nemotron-3 8B is a family of 8-billion-parameter language models from NVIDIA, introduced on 15 November 2023 with its AI foundry service on Microsoft Azure: a base model, three chat models aligned…

      Primary source Available · no news linked

Lineage

  1. Nemotron-3 8B
    Successors
    Nemotron-4 15B
  2. Nemotron-4 15B
    Predecessors
    Nemotron-3 8B
    Variants and derived
    Nemotron-4 4B Instruct · derived: distillation
  3. Nemotron-4 340B

    No recorded relations.

  4. Mistral NeMo
    Predecessors
    Mistral 7B v0.3 (Mistral AI)
    Variants and derived
    Pixtral 12B (Mistral AI) · derived: other
  5. Nemotron-4 4B Instruct
    Based on
    Nemotron-4 15B · derived: distillation
  6. Llama-3.1-Nemotron-51B-Instruct
    Based on
    Llama 3.1 (Meta) · derived: distillation
  7. Llama-3.1-Nemotron-70B-Instruct
    Based on
    Llama 3.1 (Meta) · derived: fine tune
  8. Llama-3.1-Nemotron-Nano-8B-v1
    Based on
    Llama 3.1 (Meta) · derived: fine tune
  9. Llama-3.3-Nemotron-Super-49B-v1
    Based on
    Llama 3.3 (Meta) · derived: distillation
  10. Llama-3.1-Nemotron-Ultra-253B-v1
    Based on
    Llama 3.1 (Meta) · derived: distillation
  11. Nemotron-H

    No recorded relations.

  12. Llama-3.1-Nemotron-Nano-4B-v1.1
  13. Llama-3.3-Nemotron-Super-49B-v1.5
    Based on
    Llama 3.3 (Meta) · derived: distillation
  14. Nemotron Nano 2
    Successors
    Nemotron 3 Nano
    Variants and derived
    Nemotron 3 Nano 4B · derived: distillation; Nemotron Nano 2 VL · derived: other
  15. Nemotron Nano 2 VL
    Based on
    Nemotron Nano 2 · derived: other
  16. Nemotron 3 Nano
    Predecessors
    Nemotron Nano 2
    Variants and derived
    Nemotron 3 Nano Omni · derived: other
  17. Nemotron 3 Super
  18. Nemotron 3 Nano 4B
    Based on
    Nemotron Nano 2 · derived: distillation
  19. Nemotron 3 Nano Omni
    Predecessors
    Nemotron Nano 2 VL
    Based on
    Nemotron 3 Nano · derived: other
  20. Nemotron 3 Ultra
  21. Nemotron 3.5 Lightning
    Predecessors
    Nemotron 3 Nano

Families

About the organisation

  • Description NVIDIA is a technology company headquartered in Santa Clara, California, founded in 1993 to bring 3D graphics to gaming and multimedia. It describes itself as a pioneer of accelerated computing and designs chips, systems and software for AI computing. It also develops the Nemotron family of open models and publishes their weights and training data on Hugging Face. Primary source[1] [2] [3] [4]

Company · https://www.nvidia.com · 20 models listed here

Sources

  1. [1] Contact Us: Americas Locations & Regional Offices

    Other · NVIDIA

    Provenance
    Primary
    Availability
    Active
    Last checked
  2. [2] Our History: Innovations Over the Years

    Other · NVIDIA

    Provenance
    Primary
    Availability
    Active
    Last checked
  3. [3] About Us: Company Leadership, History, Jobs, News

    Other · NVIDIA

    Provenance
    Primary
    Availability
    Active
    Last checked
  4. [4] Build Agentic AI with Multimodal Foundation Models | NVIDIA Nemotron

    Other · NVIDIA

    Provenance
    Primary
    Availability
    Active
    Last checked