Llama-3.1-Nemotron-Nano-4B-v1.1

Available · Language, Reasoning

Llama-3.1-Nemotron-Nano-4B-v1.1 is a 4B open-weights reasoning language model in NVIDIA's Llama Nemotron Nano line, fine-tuned from a Minitron-compressed Llama 3.1 8B base and also offered as an NVIDIA NIM. NVIDIA says it is small enough to run at the edge on Jetson and RTX GPUs; reasoning can be switched on or off through the system prompt. [1] [2] Primary source

Timeline of Llama-3.1-Nemotron-Nano-4B-v1.1 →

Claims and evidence

  • Released Primary source[1]Release Date; Model Version
  • Status AvailableOpen weights still downloadable from NVIDIA's Hugging Face organisation. NVIDIA has marked only its hosted API endpoint on build.nvidia.com as deprecated (see notes); no deprecation of the model weights was found. Primary source[1]model card and files
  • Successor of Llama-3.1-Nemotron-Nano-8B-v1NVIDIA's article on Nano 4B v1.1 refers to Llama 3.1 Nemotron Nano 8B v1 as the previous version. Primary source[2]Training Recipe for Llama 3.1 Nemotron Nano 4B v1.1
  • Change · Size About half the parameters of Nano 8B v1: 4 billion instead of 8 billion.Compared with Llama-3.1-Nemotron-Nano-8B-v1 Primary source[2]Advancing Reasoning SLMs for On-device Agentic AI [3]Nano
  • Change · Architecture Built on Llama-3.1-Minitron-4B-Width-Base, which NVIDIA pruned and distilled from Llama 3.1 8B, instead of directly on Llama-3.1-8B-Instruct.Compared with Llama-3.1-Nemotron-Nano-8B-v1 Primary source[1]Model Overview [4]Model Overview
  • Input text Primary source[1]Input
  • Output text Primary source[1]Output
  • Feature Reasoning modeReasoning on or off is selected through the system prompt. Primary source[1]Quick Start and Usage Recommendations [2]How to Use Llama 3.1 Nemotron Nano 4B v1.1
  • Feature Tool use Primary source[1]Model Overview [2]Function Calling with Llama 3.1 Nemotron Nano 4B v1.1
  • Open weights YesGoverned by the NVIDIA Open Model License; the Llama 3.1 Community License Agreement also applies (Built with Llama). Primary source[1]License/Terms of Use
  • Context window 131,072 tokens tokensThe model overview also gives the context length as 128K. Primary source[1]Input
  • Parameters 4 billion Primary source[2]Advancing Reasoning SLMs for On-device Agentic AI
  • Access Open-weights download, APICheckpoints on Hugging Face and an NVIDIA NIM on build.nvidia.com; NVIDIA has since marked the hosted API deprecated (no longer supported after 04/17/2026). Primary source[2]Try It Now! [5]endpoint page

Lineage

Predecessors

Successors

No known successor.

Based on

Not derived from another model.

Variants and derived

None recorded.

Siblings

None recorded.

All ancestors

All descendants

None.

Variants

No variants recorded in this record.

Related AI Radar coverage

AI Radar coverage starts in June 2026; no coverage linked yet.

All model releases from NVIDIA on AI Radar →