Llama-3.1-Nemotron-51B-Instruct

Available · Language

Llama-3.1-Nemotron-51B-Instruct is a 51-billion-parameter chat model from NVIDIA, released on 23 September 2024 and derived from Meta's Llama-3.1-70B through block-wise distillation and neural architecture search. It was offered as an NVIDIA NIM through an API and as open weights on Hugging Face; NVIDIA says it fits on a single H100 GPU. [1] [2] Primary source

Timeline of Llama-3.1-Nemotron-51B-Instruct →

Claims and evidence

  • Released Primary source[1]page date; first paragraph
  • NVIDIA-hosted API endpoint deprecated Applies to the hosted API in the NVIDIA API catalog only, not to the downloadable weights. Primary source[3]deprecation notice
  • Status AvailableWeights remain downloadable from NVIDIA's Hugging Face organisation. The NVIDIA-hosted API endpoint was deprecated as of 2025-10-10 (see milestones). Primary source[2]
  • Derived from (distillation) Llama 3.1 (Meta) Primary source[2]derived from Llama-3.1-70B-Instruct
    • Input text Primary source[2]Model Input
    • Output text Primary source[2]Model Output
    • Open weights YesNVIDIA Open Model License, with the Llama 3.1 Community License Agreement as additional information. Primary source[2]License; Quick Start
    • Parameters 51B Primary source[1]Building the model with NAS
    • Access Open-weights download, APIOffered as an NVIDIA NIM microservice through the API at ai.nvidia.com and as weights on Hugging Face. The hosted endpoint was later deprecated. Primary source[1]Simplifying inference with NVIDIA NIM [2]Quick Start

    Lineage

    Predecessors

    No known predecessor.

    Successors

    No known successor.

    Based on

    • Llama 3.1 · Meta · 23 July 2024 · derived (distillation)

    Variants and derived

    None recorded.

    Siblings

    None recorded.

    All ancestors

    All descendants

    None.

    Variants

    No variants recorded in this record.

    Related AI Radar coverage

    AI Radar coverage starts in June 2026; no coverage linked yet.

    All model releases from NVIDIA on AI Radar →