Nemotron-H

Available · Language

Nemotron-H is an NVIDIA series of hybrid Mamba-Transformer language models that replace most self-attention layers with Mamba-2 layers to lower inference cost. The 8B, 47B and 56B base checkpoints were released in April 2025 as weights for research and development use, followed in June by 8B and 47B reasoning checkpoints with a 128K context. [1] [2] [3] [4] [5] Primary source

Timeline of Nemotron-H →

Claims and evidence

  • Announced The research page announced the family and said the models would be released; the checkpoints followed in April 2025.Other sources give: [1]Internet Archive capture of 2025-03-22, Published date The earliest archived copy of the research page (captured 2025-03-22) shows Published: March 20, 2025; captures from April 2025 onward and the live page show March 21, 2025. Primary source[1]Published date
  • Released The 8B and 56B base cards give 4/14/2025 and the 47B base card gives 4/12/2025. Archived copies of the research page still said NVIDIA will be releasing the models on 2025-04-14 and said the base models had been released on 2025-04-15. Primary source[2]Release Date [3]Release Date [1]Model Release and Technical Report; Internet Archive captures of 2025-04-14 and 2025-04-15
  • Status AvailableWeights still downloadable from NVIDIA's Hugging Face organisation under the NVIDIA Internal Scientific Research and Development Model License, which limits use to research and development. Primary source[2]model card and files [1]Model Release and Technical Report
      • Input text Primary source[2]Input
      • Output text Primary source[2]Output
      • Feature MultilingualTen languages listed: English, German, Spanish, French, Italian, Korean, Portuguese, Russian, Japanese and Chinese. Primary source[2]Model Overview
      • Open weights YesReleased under the NVIDIA Internal Scientific Research and Development Model License (research and development use). Primary source[1]Model Release and Technical Report [2]License/Terms of Use [6]abstract
      • Context window 8K tokensApplies to the base checkpoints; the reasoning and instruct checkpoints state 128K. Primary source[2]Model Overview; Input [3]Model Overview; Input
      • Access Open-weights downloadCheckpoints in Hugging Face and NeMo formats on Hugging Face and NGC. Primary source[1]Model Release and Technical Report

      Lineage

      Predecessors

      No known predecessor.

      Successors

      No known successor.

      Based on

      Not derived from another model.

      Variants and derived

      None recorded.

      Siblings

      None recorded.

      All ancestors

      None.

      All descendants

      None.

      Variants

      Nemotron-H-8B-Base-8K

      • Released Primary source[2]Release Date
      • Variant Inline variant in this record. Primary source[2]

      Also known as: nvidia/Nemotron-H-8B-Base-8K, Nemotron-H-8B-Base

      Differs in:

      • Parameters: 8B Evidence not assessed [2]Model Architecture

      Nemotron-H-47B-Base-8K

      • Released Primary source[7]Release Date
      • Variant Inline variant in this record.Pruned and distilled from Nemotron-H-56B-Base-8K with the MiniPuzzle method, using 63B tokens. Primary source[7]Model Overview [1]miniPuzzle compression section

      Also known as: nvidia/Nemotron-H-47B-Base-8K, Nemotron-H-47B-Base

      Differs in:

      • Parameters: 47B Evidence not assessed [7]Model Architecture

      Nemotron-H-56B-Base-8K

      • Released Primary source[3]Release Date
      • Variant Inline variant in this record. Primary source[3]

      Also known as: nvidia/Nemotron-H-56B-Base-8K, Nemotron-H-56B-Base

      Differs in:

      • Parameters: 56B Evidence not assessed [3]Model Architecture

      Nemotron-H-8B-Reasoning-128K

      • Released Primary source[4]Release Date
      • Variant Inline variant in this record.Post-trained reasoning model based on Nemotron-H-8B-Base-8K; BF16 checkpoint, with an FP8 checkpoint in a separate repository. Primary source[4]Model Overview

      Also known as: nvidia/Nemotron-H-8B-Reasoning-128K, Nemotron-H-8B-Reasoning-128K-FP8

      Differs in:

      • Context window: 128K tokens Evidence not assessed [4]Input
      • Parameters: 8B Evidence not assessed [4]Model Architecture

      Nemotron-H-47B-Reasoning-128K

      • Released Primary source[5]Release Date
      • Variant Inline variant in this record.Post-trained reasoning model based on Nemotron-H-47B-Base-8K; BF16 checkpoint, with an FP8 checkpoint in a separate repository. Primary source[5]Model Overview

      Also known as: nvidia/Nemotron-H-47B-Reasoning-128K, Nemotron-H-47B-Reasoning-128K-FP8

      Differs in:

      • Context window: 128K tokens Evidence not assessed [5]Input
      • Parameters: 47B Evidence not assessed [5]Model Architecture

      Nemotron-H-4B-Base-8K

      • Released Primary source[8]Release Date
      • Variant Inline variant in this record.Pruned and distilled from Nemotron-H-8B-Base-8K using 380B tokens. The card gives the Hugging Face release as 10/23/2025, although Hugging Face metadata shows the repository was created on 2025-03-20. Primary source[8]Model Overview

      Also known as: nvidia/Nemotron-H-4B-Base-8K

      Differs in:

      • Parameters: 4B Evidence not assessed [8]Model Architecture

      Nemotron-H-4B-Instruct-128K

      • Released Primary source[9]Release Date
      • Variant Inline variant in this record.Aligned version of Nemotron-H-4B-Base-8K for chat, instruction following and tool calling. The card gives the Hugging Face release as 10/23/2025, although Hugging Face metadata shows the repository was created on 2025-04-15. Primary source[9]Model Overview

      Also known as: nvidia/Nemotron-H-4B-Instruct-128K

      Differs in:

      • Context window: 128K tokens Evidence not assessed [9]Input
      • Parameters: 4B Evidence not assessed [9]Model Architecture

      Related AI Radar coverage

      AI Radar coverage starts in June 2026; no coverage linked yet.

      All model releases from NVIDIA on AI Radar →