Llama-3.1-Nemotron-70B-Instruct

Available · Language

Llama-3.1-Nemotron-70B-Instruct is a language model from NVIDIA, trained from Meta's Llama-3.1-70B-Instruct with reinforcement learning from human feedback and NVIDIA's own reward model to make responses more helpful. Presented in October 2024, it was published as open weights and offered through a hosted API; NVIDIA calls it a demonstration of its helpfulness techniques. [1] [2] Primary source

Timeline of Llama-3.1-Nemotron-70B-Instruct →

Claims and evidence

  • Released No NVIDIA page states a day. The HelpSteer2-Preference paper v1 of 2024-10-02 announces only the reward model release; the reward-model blog (dated 2024-10-03, updated 10/21/2024) presents the Instruct model and links to it, and the v2 abstract mentions its open release. Primary source[2]Leading large language model; Getting started; update note of 10/21/2024 [3]Submission history; abstract v1 and v2
  • NVIDIA-hosted API endpoint deprecated Applies to the hosted API in the NVIDIA API catalog only, not to the downloadable weights. Primary source[4]deprecation notice
  • Status AvailableWeights remain downloadable from NVIDIA's Hugging Face organisation in NeMo and Transformers formats. The NVIDIA-hosted API endpoint was deprecated as of 2025-10-10 (see milestones). Primary source[1] [5]
  • Derived from (fine tune) Llama 3.1 (Meta) Primary source[1]initial policy: Llama-3.1-70B-Instruct
    • Input text Primary source[1]Input
    • Output text Primary source[1]Output
    • Open weights YesThe current cards name the NVIDIA Open Model License, with the Llama 3.1 Community License Agreement as additional information; the reward-model blog described the model as coming with the Llama 3.1 license. Primary source[1]License; Usage [5]License [3]Abstract (v2)
    • Context window 128k tokensInput limit as stated on the card (max of 128k tokens); output is given as a maximum of 4k tokens. Primary source[1]Input: Other Properties Related to Input
    • Parameters 70BAs given in NVIDIA's model name; the card states no separate parameter count. Primary source[1]
    • Access Open-weights download, APIHosted inference with an OpenAI-compatible API at build.nvidia.com, and weights on Hugging Face. The hosted endpoint was later deprecated. Primary source[1]Description; Usage [4]

    Lineage

    Predecessors

    No known predecessor.

    Successors

    No known successor.

    Based on

    • Llama 3.1 · Meta · 23 July 2024 · derived (fine tune)

    Variants and derived

    None recorded.

    Siblings

    None recorded.

    All ancestors

    All descendants

    None.

    Variants

    No variants recorded in this record.

    Related AI Radar coverage

    AI Radar coverage starts in June 2026; no coverage linked yet.

    All model releases from NVIDIA on AI Radar →