Llama-3.3-Nemotron-Super-49B-v1.5

Available · Language, Reasoning

Llama-3.3-Nemotron-Super-49B-v1.5 is an open-weights NVIDIA reasoning language model of about 49 billion parameters, derived from Meta's Llama-3.3-70B-Instruct and released on Hugging Face and as a preview API on build.nvidia.com. NVIDIA calls it a significantly upgraded version of Super v1, post-trained for reasoning, chat and agentic tasks; reasoning is on by default and can be switched off. [1] [2] Primary source

Timeline of Llama-3.3-Nemotron-Super-49B-v1.5 →

Claims and evidence

  • Released The developer blog header shows Jul 29, 2025, the date of an update with leaderboard information; the post says it first ran on July 25, the Friday on which the model was released. Primary source[1]Release Date; Model Version [2]introduction (latest version released Friday) and closing line (originally ran July 25)
  • Status AvailableNVIDIA has no deprecation page for its open models; the weights were still downloadable from NVIDIA's Hugging Face organisation on 2026-10-01. Primary source[1]
  • Successor of Llama-3.3-Nemotron-Super-49B-v1The model card calls it a significantly upgraded version of Llama-3.3-Nemotron-Super-49B-v1; recorded as successor-of rather than revision-of because the version in the name changed to v1.5. Primary source[1]Model Overview
  • Derived from (distillation) Llama 3.3 (Meta) Primary source[1]derived from Llama-3.3-70B-Instruct
  • Change · Reasoning Further post-trained on additional reasoning data; NVIDIA says this improves math, science, coding, function calling, instruction following and chat.Compared with Llama-3.3-Nemotron-Super-49B-v1 Primary source[2]
  • Change · Tool use NVIDIA trained tool calling with iterative DPO stages, and the model card provides a tool-call parser for serving the model with vLLM.Compared with Llama-3.3-Nemotron-Super-49B-v1 Primary source[1]
  • Input text Primary source[1]Input
  • Output text Primary source[1]Output
  • Feature Reasoning modeReasoning on by default; /no_think in the system prompt switches it off. Primary source[1]Quick Start and Usage Recommendations
  • Feature Function calling Primary source[1]Running a vLLM Server with Tool-call Support
  • Open weights YesWeights published on Hugging Face under the NVIDIA Open Model License, with the Llama 3.3 Community License Agreement as additional information. Primary source[1]License/Terms of Use; Release Date
  • Context window 128K tokensThe card also gives the figure as up to 131,072 tokens. Primary source[1]Model Overview; Input
  • Parameters c. 49BFigure as given in the official model name; the card text gives no separate count. Primary source[1]model name
  • Access Open-weights download, APIReleased on Hugging Face and as a preview API on build.nvidia.com; the blog said an NVIDIA NIM microservice would follow soon. Primary source[1]Release Date; Quick Start and Usage Recommendations [2]Get started with Llama Nemotron Super v1.5

Lineage

Predecessors

Successors

Based on

  • Llama 3.3 · Meta · 6 December 2024 · derived (distillation)

Variants and derived

None recorded.

Siblings

None recorded.

All ancestors

All descendants

Variants

No variants recorded in this record.

Related AI Radar coverage

AI Radar coverage starts in June 2026; no coverage linked yet.

All model releases from NVIDIA on AI Radar →