Claims and evidence
- Announced Announced with the Nemotron 3 family, then described as about 500B parameters with up to 50B active and expected in the first half of 2026. Jensen Huang presented it again in the GTC Taipei keynote at COMPUTEX (live blog keynote recap dated May 31, 2026, 8:00 p.m. PT). Primary source[4]
- Released Primary source[1]Thursday, June 4, 6:00 a.m. PT [2]Model Summary: Release Date; Release Date [3]Published
- Status AvailableNVIDIA has no deprecation page for its open models; the weights were still downloadable from the NVIDIA organisation on Hugging Face on 2026-10-01. Primary source[2]
- Successor of Llama-3.1-Nemotron-Ultra-253B-v1 EditorialEditorial link along the Ultra tier: NVIDIA does not name a predecessor; Llama-3.1-Nemotron-Ultra-253B-v1 was the previous Ultra-tier release. Primary source[3]
- Change · Architecture Uses a hybrid Mamba-attention LatentMoE architecture with multi-token prediction layers, pre-trained in NVFP4.Compared with Llama-3.1-Nemotron-Ultra-253B-v1 Primary source[3]
- Change · Size 550B total parameters, of which 55B are active per token.Compared with Llama-3.1-Nemotron-Ultra-253B-v1 Primary source[2]Model Summary: Total Parameters
- Change · Context length Supports a context length of up to 1M tokens.Compared with Llama-3.1-Nemotron-Ultra-253B-v1 Primary source[2]Model Summary: Context Length
- Change · Reasoning Supports inference-time control of the reasoning budget.Compared with Llama-3.1-Nemotron-Ultra-253B-v1 Primary source[3]Key Features
- Input text Primary source[2]Input
- Output text Primary source[2]Output
- Feature Reasoning modeReasoning can be switched on or off through the chat template; the research page also lists inference-time reasoning budget control. Primary source[2]Model Summary: Reasoning Mode [3]Key Features
- Feature Tool use Primary source[2]Model Summary: Best For
- Feature MultilingualEnglish, French, Spanish, Italian, German, Japanese, Hindi, Korean, Brazilian Portuguese and Chinese. Primary source[2]Model Summary: Supported Languages
- Open weights YesReleased under the OpenMDW License Agreement, version 1.1. Primary source[3]Open Source [2]License/Terms of Use
- Context window 1M tokensStated as up to 1M tokens. Primary source[2]Model Summary: Context Length [3]Key Highlights
- Parameters 550B (55B active) Primary source[2]Model Summary: Total Parameters [3]
- Access Open-weights download, API, cloud partnerHugging Face, ModelScope, OpenRouter, build.nvidia.com as NVIDIA NIM microservices and partner platforms. Primary source[1]Thursday, June 4, 6:00 a.m. PT: Open and Customizable, Deployable Anywhere
Lineage
Predecessors
- Llama-3.1-Nemotron-Ultra-253B-v1 · 7 April 2025
Successors
No known successor.
Based on
Not derived from another model.
Variants and derived
None recorded.
Siblings
None recorded.
All ancestors
- LLaMA · Meta · 24 February 2023
- Llama 2 · Meta · 18 July 2023
- Meta Llama 3 · Meta · 18 April 2024
- Llama 3.1 · Meta · 23 July 2024
- Llama-3.1-Nemotron-Ultra-253B-v1 · 7 April 2025
All descendants
None.
Variants
Nemotron 3 Ultra 550B-A55B
Same dates as Nemotron 3 Ultra.
- Variant Inline variant in this record.Post-trained model, published in BF16 and as an NVFP4 quantized checkpoint. Primary source[3]Open Source: Checkpoints [2]
Also known as: NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16, NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4
Nemotron 3 Ultra 550B-A55B Base
Same dates as Nemotron 3 Ultra.
- Variant Inline variant in this record.Pre-trained base checkpoint. Primary source[3]Open Source: Checkpoints [2]Training Methodology: Stage 1
Also known as: NVIDIA-Nemotron-3-Ultra-550B-A55B-Base-BF16
Related AI Radar coverage
- How Full-Stack NIM Optimizations Deliver 2.5x More Users on Nemotron 3 UltraNVIDIA · Developer ·