Claims and evidence
- Announced Primary source[1]Submission history, v1
- ReleasedRelease date unknown
- Status Research onlyDescribed in a technical report only. As of 2026-10-01 no NVIDIA weight release (Hugging Face or NGC) or NVIDIA-hosted access for this model was found. Primary source[1]
- Successor of Nemotron-3 8B EditorialEditorial link along the Nemotron generation numbers: the report introduces Nemotron-4 without naming Nemotron-3 8B as its predecessor. Primary source[1]
- Change · Size 15 billion parameters, up from 8 billion for Nemotron-3 8B.Compared with Nemotron-3 8B Primary source[1] [2]
- Change · Training data Pre-trained on 8 trillion tokens covering 43 programming languages, against 3.8 trillion tokens and 37 programming languages for the Nemotron-3 8B base model.Compared with Nemotron-3 8B Primary source[1] [2]
- Change · Architecture Uses grouped-query attention, rotary position embeddings and squared ReLU activations; NVIDIA's Nemotron-3 8B model cards describe a GPT-3 style network.Compared with Nemotron-3 8B Primary source[1] [2]
- Feature MultilingualPre-training data covers 53 natural languages besides English and 43 programming languages. Primary source[1]Abstract; section 2, Data
- Context window 4096 tokensGiven as the sequence length in the report's hyper-parameter table. Primary source[1]Table 1, sequence length
- Parameters 15 billionThe report splits this into 3.2 billion embedding and 12.5 billion non-embedding parameters. Primary source[1]Abstract; section 2
Lineage
Predecessors
- Nemotron-3 8B · 15 November 2023
Successors
No known successor.
Based on
Not derived from another model.
Variants and derived
- Nemotron-4 4B Instruct · 20 August 2024 · derived (distillation)
Siblings
None recorded.
All ancestors
- Nemotron-3 8B · 15 November 2023
All descendants
- Nemotron-4 4B Instruct · 20 August 2024
Variants
No variants recorded in this record.
Related AI Radar coverage
AI Radar coverage starts in June 2026; no coverage linked yet.