Llama-3.1-Nemotron-Ultra-253B-v1

Available · Language, Reasoning · Milestone

Llama-3.1-Nemotron-Ultra-253B-v1 is a 253B open-weights reasoning language model from NVIDIA, the Ultra size of the Llama Nemotron family, released in April 2025 and also offered as a hosted preview API. NVIDIA derived it from Meta's Llama-3.1-405B-Instruct with neural architecture search, distillation and continued pretraining, and says it fits on a single 8xH100 node. [1] [2] [3] Primary source

Timeline of Llama-3.1-Nemotron-Ultra-253B-v1 →

Claims and evidence

  • Announced The GTC press release introduces the Ultra size as forthcoming; its availability section lists only Nano and Super. Primary source[2]NVIDIA Post-Training Boosts Accuracy and Reliability for Enterprise Reasoning
  • Released Primary source[1]Release Date; Model Version
  • Status AvailableOpen weights still downloadable from NVIDIA's Hugging Face organisation. NVIDIA has marked only its hosted API endpoint on build.nvidia.com as deprecated (see notes); no deprecation of the model weights was found. Primary source[1]model card and files
  • Derived from (distillation) Llama 3.1 (Meta) Primary source[1]derived from Llama-3.1-405B-Instruct
    • Input text Primary source[1]Input
    • Output text Primary source[1]Output
    • Feature Reasoning modeReasoning on or off is selected through the system prompt. Primary source[1]Quick Start and Usage Recommendations [4]abstract
    • Feature Tool use Primary source[1]Model Overview
    • Open weights YesGoverned by the NVIDIA Open Model License; the Llama 3.1 Community License Agreement also applies (Built with Llama). Primary source[1]License/Terms of Use
    • Context window 131,072 tokens tokensThe model overview also gives the context length as 128K tokens. Primary source[1]Input
    • Parameters 253B Primary source[1]Model Architecture [3]Ultra
    • Access Open-weights download, APIHosted preview API on build.nvidia.com; NVIDIA has since marked the hosted API deprecated (deprecation date 04/22/2026), although the id is still listed in its public model list on 2026-10-02. Primary source[1]Quick Start and Usage Recommendations [3]Get started with NVIDIA Llama Nemotron models [5]endpoint page

    Lineage

    Predecessors

    No known predecessor.

    Successors

    Based on

    • Llama 3.1 · Meta · 23 July 2024 · derived (distillation)

    Variants and derived

    None recorded.

    Siblings

    None recorded.

    All ancestors

    All descendants

    Variants

    Llama-3.1-Nemotron-Ultra-253B-CPT-v1

    • Released Primary source[6]Release Date; Model Version
    • Variant Inline variant in this record.Base checkpoint after knowledge distillation and continued pretraining, before the reasoning post-training; NVIDIA says it can be used as a base model. Primary source[6]Model Overview

    Also known as: nvidia/Llama-3_1-Nemotron-Ultra-253B-CPT-v1

    Related AI Radar coverage

    AI Radar coverage starts in June 2026; no coverage linked yet.

    All model releases from NVIDIA on AI Radar →