Llama-3.3-Nemotron-Super-49B-v1

Available · Language, Reasoning · Milestone

Llama-3.3-Nemotron-Super-49B-v1 is an open-weights reasoning language model from NVIDIA, released at GTC 2025 as the Super size of the Llama Nemotron family and also offered as a hosted API. NVIDIA reduced Meta's Llama-3.3-70B-Instruct to 49B parameters with neural architecture search and distillation, targeting a single H100 GPU; reasoning is switched on or off in the system prompt. [1] [2] [3] Primary source

Timeline of Llama-3.3-Nemotron-Super-49B-v1 →

Claims and evidence

  • Released Primary source[1]Release Date; Model Version [2]dateline; Availability
  • Status AvailableOpen weights still downloadable from NVIDIA's Hugging Face organisation. NVIDIA has marked only its hosted API endpoint on build.nvidia.com as deprecated (see notes); no deprecation of the model weights was found. Primary source[1]model card and files
  • Derived from (distillation) Llama 3.3 (Meta) Primary source[1]derived from Llama-3.3-70B-Instruct
    • Input text Primary source[1]Input
    • Output text Primary source[1]Output
    • Feature Reasoning modeReasoning on or off is selected through the system prompt. Primary source[1]Quick Start and Usage Recommendations [3]Overview of test-time scaling
    • Feature Tool use Primary source[1]Model Overview
    • Open weights YesGoverned by the NVIDIA Open Model License; the Llama 3.3 Community License Agreement also applies (Built with Llama). Primary source[1]License/Terms of Use
    • Context window 131,072 tokens tokensThe model overview also gives the context length as 128K tokens. Primary source[1]Input
    • Parameters 49B Primary source[3]Super [4]abstract
    • Access Open-weights download, APIHosted preview API and NIM microservice on build.nvidia.com at release; NVIDIA has since marked the hosted API deprecated (no longer supported after 08/25/2026). Primary source[2]Availability [1]Quick Start and Usage Recommendations [5]endpoint page

    Lineage

    Predecessors

    No known predecessor.

    Successors

    Based on

    • Llama 3.3 · Meta · 6 December 2024 · derived (distillation)

    Variants and derived

    None recorded.

    Siblings

    None recorded.

    All ancestors

    All descendants

    Variants

    No variants recorded in this record.

    Related AI Radar coverage

    AI Radar coverage starts in June 2026; no coverage linked yet.

    All model releases from NVIDIA on AI Radar →