NVIDIA models
NVIDIA is a technology company headquartered in Santa Clara, California, founded in 1993 to bring 3D graphics to gaming and multimedia. It describes itself as a pioneer of accelerated computing and designs chips, systems and software for AI computing. It also develops the Nemotron family of open models and publishes their weights and training data on Hugging Face.
Chronology
-
2026
5 models-
ReleasedNemotron Lightning · Language, Reasoning Nemotron 3.5 Lightning
Nemotron 3.5 Lightning is an open-weights reasoning language model with 30B total and 3B active parameters that NVIDIA added to the Nemotron 3 family after Nemotron 3 Nano. NVIDIA aims it at…
- Includes multi-token prediction layers from training onward and ships with DFlash and DSpark draft models for speculative decoding.
- Pre-trained on more than 20 trillion tokens with an NVFP4 recipe; the card gives May 2026 as the post-training data cutoff.
Primary source Available · 1 news item -
ReleasedNemotron Ultra · Language, Reasoning · Milestone Nemotron 3 Ultra
Nemotron 3 Ultra is an open-weights reasoning language model and the Ultra size of NVIDIA's Nemotron 3 family, with 550B total and 55B active parameters and a context length of up to 1M tokens…
- Uses a hybrid Mamba-attention LatentMoE architecture with multi-token prediction layers, pre-trained in NVFP4.
- 550B total parameters, of which 55B are active per token.
Primary source Available · 1 news item -
ReleasedNemotron Nano · Language, Multimodal, Reasoning Nemotron 3 Nano Omni
Nemotron 3 Nano Omni is an open-weights multimodal reasoning model in NVIDIA's Nemotron 3 family that takes video, audio, images and text and produces text. It pairs the Nemotron 3 Nano 30B-A3B…
- Accepts audio input through a Parakeet speech encoder, alongside video, image and text input.
- Uses the Nemotron 3 Nano 30B-A3B hybrid mixture-of-experts language model as its backbone, with the C-RADIOv4-H vision encoder.
Primary source Available · no news linked -
ReleasedNemotron Nano · Language, Reasoning Nemotron 3 Nano 4B
Nemotron 3 Nano 4B is a small open-weights text model in NVIDIA's Nemotron 3 family that handles reasoning and non-reasoning tasks. NVIDIA calls it its first model optimised for on-device use and…
- Compressed from the 9B Nemotron Nano 2 model to about 4B parameters.
- Pruning cut depth from 56 to 42 layers and also reduced Mamba heads, feed-forward width and embedding size.
Primary source Available · no news linked -
ReleasedNemotron Super · Language, Reasoning · Milestone Nemotron 3 Super
Nemotron 3 Super is an open-weights reasoning language model in NVIDIA's Nemotron 3 family, with 120B total and 12B active parameters, a hybrid Mamba-Transformer mixture-of-experts design and a…
- Uses a hybrid Mamba-Transformer latent mixture-of-experts design with multi-token prediction, activating 12B of its 120B parameters per token.
- Supports a context window of up to 1M tokens.
Primary source Available · no news linked
-
-
2025
9 models-
ReleasedNemotron Nano · Language, Reasoning · Milestone Nemotron 3 Nano
Nemotron 3 Nano is the first released model of NVIDIA's Nemotron 3 open model family: a hybrid Mamba-Transformer mixture-of-experts reasoning model with 31.6 billion total and 3.2 billion active…
- Moves to a mixture-of-experts hybrid Mamba-Transformer design that activates less than half as many parameters per forward pass as Nemotron 2 Nano.
- Context window grows from 128K tokens to up to 1M tokens.
Primary source Available · no news linked -
ReleasedNemotron Nano · Language, Multimodal, Reasoning Nemotron Nano 2 VL
Nemotron Nano 2 VL is a 12B open-weights vision-language model from NVIDIA that takes text, images and video and produces text, built on the Nemotron Nano 2 12B reasoning model with a RADIO vision…
- Adds image and video input to the text-only Nemotron Nano 2 language model it is built on.
- Introduces Efficient Video Sampling, which drops video patches that stay unchanged over time to cut token counts for long videos.
Primary source Available · no news linked -
ReleasedNemotron Nano · Language, Reasoning · Milestone Nemotron Nano 2
Nemotron Nano 2 is a family of open-weights hybrid Mamba-Transformer reasoning models that NVIDIA released on Hugging Face with 9B and 12B checkpoints and a 128K-token context. The 9B reasoning…
- Hybrid Mamba-2 and Transformer design trained from scratch by NVIDIA, whereas the v1 Nano model was derived from a Llama model.
- Adds a thinking budget that can be set at inference time, next to switching reasoning on or off.
Primary source Available · no news linked -
ReleasedNemotron Super · Language, Reasoning Llama-3.3-Nemotron-Super-49B-v1.5
Llama-3.3-Nemotron-Super-49B-v1.5 is an open-weights NVIDIA reasoning language model of about 49 billion parameters, derived from Meta's Llama-3.3-70B-Instruct and released on Hugging Face and as a…
- Further post-trained on additional reasoning data; NVIDIA says this improves math, science, coding, function calling, instruction following and chat.
- NVIDIA trained tool calling with iterative DPO stages, and the model card provides a tool-call parser for serving the model with vLLM.
Primary source Available · no news linked -
ReleasedNemotron Nano · Language, Reasoning Llama-3.1-Nemotron-Nano-4B-v1.1
Llama-3.1-Nemotron-Nano-4B-v1.1 is a 4B open-weights reasoning language model in NVIDIA's Llama Nemotron Nano line, fine-tuned from a Minitron-compressed Llama 3.1 8B base and also offered as an…
- About half the parameters of Nano 8B v1: 4 billion instead of 8 billion.
- Built on Llama-3.1-Minitron-4B-Width-Base, which NVIDIA pruned and distilled from Llama 3.1 8B, instead of directly on Llama-3.1-8B-Instruct.
Primary source Available · no news linked -
ReleasedNemotron · Language Nemotron-H
Nemotron-H is an NVIDIA series of hybrid Mamba-Transformer language models that replace most self-attention layers with Mamba-2 layers to lower inference cost. The 8B, 47B and 56B base checkpoints…
Primary source Available · no news linked -
ReleasedNemotron Ultra · Language, Reasoning · Milestone Llama-3.1-Nemotron-Ultra-253B-v1
Llama-3.1-Nemotron-Ultra-253B-v1 is a 253B open-weights reasoning language model from NVIDIA, the Ultra size of the Llama Nemotron family, released in April 2025 and also offered as a hosted preview…
Primary source Available · no news linked -
ReleasedNemotron Nano · Language, Reasoning · Milestone Llama-3.1-Nemotron-Nano-8B-v1
Llama-3.1-Nemotron-Nano-8B-v1 is an 8B open-weights reasoning language model from NVIDIA, derived from Meta's Llama-3.1-8B-Instruct and released at GTC 2025 as the Nano size of the Llama Nemotron…
Primary source Available · no news linked -
ReleasedNemotron Super · Language, Reasoning · Milestone Llama-3.3-Nemotron-Super-49B-v1
Llama-3.3-Nemotron-Super-49B-v1 is an open-weights reasoning language model from NVIDIA, released at GTC 2025 as the Super size of the Llama Nemotron family and also offered as a hosted API. NVIDIA…
Primary source Available · no news linked
-
-
2024
6 models-
Day not givenNemotron · Language Llama-3.1-Nemotron-70B-Instruct
Llama-3.1-Nemotron-70B-Instruct is a language model from NVIDIA, trained from Meta's Llama-3.1-70B-Instruct with reinforcement learning from human feedback and NVIDIA's own reward model to make…
Primary source Available · no news linked -
ReleasedNemotron · Language Llama-3.1-Nemotron-51B-Instruct
Llama-3.1-Nemotron-51B-Instruct is a 51-billion-parameter chat model from NVIDIA, released on 23 September 2024 and derived from Meta's Llama-3.1-70B through block-wise distillation and neural…
Primary source Available · no news linked -
ReleasedNemotron · Language Nemotron-4 4B Instruct
Nemotron-4 4B Instruct is a small language model from NVIDIA, announced on 20 August 2024 for NVIDIA ACE, its game-character technology, as an NVIDIA NIM for cloud and on-device use. NVIDIA says it…
- About 4 billion parameters by name, distilled from the 15-billion-parameter Nemotron-4 15B.
- Pruned, distilled and INT4-quantised by NVIDIA for on-device inference using about 2 GB of VRAM.
Primary source Available · no news linked -
ReleasedMistral · Language Mistral NeMo
Mistral NeMo is a 12B-parameter language model that Mistral AI built with NVIDIA, released with base and instruct weights under Apache 2.0 and offered on la Plateforme and as an NVIDIA NIM. It has a…
- Has 12B parameters, up from 7B in Mistral 7B, while keeping a standard architecture that Mistral says lets it replace Mistral 7B as a drop-in.
- Uses a new tokenizer, Tekken, based on Tiktoken, in place of the SentencePiece tokenizer of earlier Mistral models.
-
ReleasedNemotron · Language Nemotron-4 340B
Nemotron-4 340B is a family of language models from NVIDIA (base, instruct and reward), released on 14 June 2024 as open weights on Hugging Face and the NGC catalog under the NVIDIA Open Model…
- 340 billion parameters, against 15 billion for Nemotron-4 15B, with a similar architecture.
- Uses the same data blend as Nemotron-4 15B but trains on 9 trillion tokens: the first 8 trillion plus 1 trillion tokens of continued pre-training.
Primary source Available · no news linked -
Announced, release date unknownNemotron · Language · Milestone Nemotron-4 15B
Nemotron-4 15B is a 15-billion-parameter multilingual language model from NVIDIA, described in a technical report posted on arXiv on 26 February 2024. It was trained on 8 trillion tokens of English…
- 15 billion parameters, up from 8 billion for Nemotron-3 8B.
- Pre-trained on 8 trillion tokens covering 43 programming languages, against 3.8 trillion tokens and 37 programming languages for the Nemotron-3 8B base model.
Primary source Research only · no news linked
-
-
2023
1 model-
ReleasedNemotron · Language · Milestone Nemotron-3 8B
Nemotron-3 8B is a family of 8-billion-parameter language models from NVIDIA, introduced on 15 November 2023 with its AI foundry service on Microsoft Azure: a base model, three chat models aligned…
Primary source Available · no news linked
-
Lineage
-
Nemotron-3 8B
- Successors
- Nemotron-4 15B
-
Nemotron-4 15B
- Predecessors
- Nemotron-3 8B
- Variants and derived
- Nemotron-4 4B Instruct · derived: distillation
-
Nemotron-4 340B
No recorded relations.
-
Mistral NeMo
- Predecessors
- Mistral 7B v0.3 (Mistral AI)
- Variants and derived
- Pixtral 12B (Mistral AI) · derived: other
-
Nemotron-4 4B Instruct
- Based on
- Nemotron-4 15B · derived: distillation
-
Llama-3.1-Nemotron-51B-Instruct
- Based on
- Llama 3.1 (Meta) · derived: distillation
-
Llama-3.1-Nemotron-70B-Instruct
- Based on
- Llama 3.1 (Meta) · derived: fine tune
-
Llama-3.1-Nemotron-Nano-8B-v1
- Successors
- Llama-3.1-Nemotron-Nano-4B-v1.1; Nemotron Nano 2
- Based on
- Llama 3.1 (Meta) · derived: fine tune
-
Llama-3.3-Nemotron-Super-49B-v1
- Successors
- Llama-3.3-Nemotron-Super-49B-v1.5
- Based on
- Llama 3.3 (Meta) · derived: distillation
-
Llama-3.1-Nemotron-Ultra-253B-v1
- Successors
- Nemotron 3 Ultra
- Based on
- Llama 3.1 (Meta) · derived: distillation
-
Nemotron-H
No recorded relations.
-
Llama-3.1-Nemotron-Nano-4B-v1.1
- Predecessors
- Llama-3.1-Nemotron-Nano-8B-v1
-
Llama-3.3-Nemotron-Super-49B-v1.5
- Predecessors
- Llama-3.3-Nemotron-Super-49B-v1
- Successors
- Nemotron 3 Super
- Based on
- Llama 3.3 (Meta) · derived: distillation
-
Nemotron Nano 2
- Predecessors
- Llama-3.1-Nemotron-Nano-8B-v1
- Successors
- Nemotron 3 Nano
- Variants and derived
- Nemotron 3 Nano 4B · derived: distillation; Nemotron Nano 2 VL · derived: other
-
Nemotron Nano 2 VL
- Successors
- Nemotron 3 Nano Omni
- Based on
- Nemotron Nano 2 · derived: other
-
Nemotron 3 Nano
- Predecessors
- Nemotron Nano 2
- Successors
- Nemotron 3.5 Lightning
- Variants and derived
- Nemotron 3 Nano Omni · derived: other
-
Nemotron 3 Super
- Predecessors
- Llama-3.3-Nemotron-Super-49B-v1.5
-
Nemotron 3 Nano 4B
- Based on
- Nemotron Nano 2 · derived: distillation
-
Nemotron 3 Nano Omni
- Predecessors
- Nemotron Nano 2 VL
- Based on
- Nemotron 3 Nano · derived: other
-
Nemotron 3 Ultra
- Predecessors
- Llama-3.1-Nemotron-Ultra-253B-v1
-
Nemotron 3.5 Lightning
- Predecessors
- Nemotron 3 Nano
Families
-
Nemotron
Nemotron-3 8B · Nemotron-4 15B · Nemotron-4 340B · Nemotron-4 4B Instruct · Llama-3.1-Nemotron-51B-Instruct · Llama-3.1-Nemotron-70B-Instruct · Nemotron-H
-
Nemotron Lightning
-
Nemotron Nano
Llama-3.1-Nemotron-Nano-8B-v1 · Llama-3.1-Nemotron-Nano-4B-v1.1 · Nemotron Nano 2 · Nemotron Nano 2 VL · Nemotron 3 Nano · Nemotron 3 Nano 4B · Nemotron 3 Nano Omni
-
Nemotron Super
Llama-3.3-Nemotron-Super-49B-v1 · Llama-3.3-Nemotron-Super-49B-v1.5 · Nemotron 3 Super
-
Nemotron Ultra
-
About the organisation
- Description NVIDIA is a technology company headquartered in Santa Clara, California, founded in 1993 to bring 3D graphics to gaming and multimedia. It describes itself as a pioneer of accelerated computing and designs chips, systems and software for AI computing. It also develops the Nemotron family of open models and publishes their weights and training data on Hugging Face. Primary source[1] [2] [3] [4]
Company · https://www.nvidia.com · 20 models listed here
Sources
-
[1] Contact Us: Americas Locations & Regional Offices
Other · NVIDIA
- Provenance
- Primary
- Availability
- Active
- Last checked
Original ↗ https://www.nvidia.com/en-us/contact/ Lists NVIDIA Corporate at 2788 San Tomas Expressway, Santa Clara, CA 95051. -
[2] Our History: Innovations Over the Years
Other · NVIDIA
- Provenance
- Primary
- Availability
- Active
- Last checked
Original ↗ https://www.nvidia.com/en-us/about-nvidia/corporate-timeline/ Corporate timeline; gives the founding date (April 5, 1993) and the three founders, with a vision to bring 3D graphics to the gaming and multimedia markets. Living page. -
[3] About Us: Company Leadership, History, Jobs, News
Other · NVIDIA
- Provenance
- Primary
- Availability
- Active
- Last checked
Original ↗ https://www.nvidia.com/en-us/about-nvidia/ Company overview page; states NVIDIA pioneered accelerated computing and engineers chips, systems and software for AI. Living page, content as seen on 2026-10-01. The HTML title ends with the site suffix | NVIDIA. -
[4] Build Agentic AI with Multimodal Foundation Models | NVIDIA Nemotron
Other · NVIDIA
- Provenance
- Primary
- Availability
- Active
- Last checked
Original ↗ https://www.nvidia.com/en-us/ai-data-science/foundation-models/nemotron/ Product page for the Nemotron model family (page heading: NVIDIA Nemotron); states models and training data are published openly on Hugging Face. Living page, content as seen on 2026-10-01.