Nemotron-4 4B Instruct

Available · Language

Nemotron-4 4B Instruct is a small language model from NVIDIA, announced on 20 August 2024 for NVIDIA ACE, its game-character technology, as an NVIDIA NIM for cloud and on-device use. NVIDIA says it was distilled from Nemotron-4 15B and tuned for role-play, retrieval-augmented generation and function calling; weights followed on Hugging Face as Nemotron-Mini-4B-Instruct. [1] [2] Primary source

Timeline of Nemotron-4 4B Instruct →

Claims and evidence

  • Released Announced at Gamescom 2024; the blog of that day says the model is available as an NVIDIA NIM. The Hugging Face weights carry no stated date. Primary source[1]page date; first section
  • NVIDIA-hosted API endpoint deprecated The notice, as seen on 2026-10-01, says the API will be deprecated on 08/25/2026. Applies to the hosted API only, not to the downloadable weights. Primary source[3]deprecation notice
  • Status AvailableWeights remain downloadable from NVIDIA's Hugging Face organisation as Nemotron-Mini-4B-Instruct. The NVIDIA API catalog set a deprecation date for the hosted endpoint (see milestones). Primary source[2]
  • Derived from (distillation) Nemotron-4 15BThe card says it is a fine-tuned version of Minitron-4B-Base, which was pruned and distilled from Nemotron-4 15B. Primary source[1]NVIDIA ACE introduces an SLM purpose-built for roleplaying [2]Model Overview
  • Change · Size About 4 billion parameters by name, distilled from the 15-billion-parameter Nemotron-4 15B.Compared with Nemotron-4 15B Primary source[1] [2]
  • Change · Efficiency Pruned, distilled and INT4-quantised by NVIDIA for on-device inference using about 2 GB of VRAM.Compared with Nemotron-4 15B Primary source[1]
  • Change · Tool use Instruction-tuned for function calling, role-play and retrieval-augmented generation; Nemotron-4 15B was a base model described in a report only.Compared with Nemotron-4 15B Primary source[1] [4]
  • Feature Function calling Primary source[1]NVIDIA ACE introduces an SLM purpose-built for roleplaying [2]Model Overview; Prompt Format, Tool use
  • Open weights YesPublished on Hugging Face under the NVIDIA Community Model License. Primary source[2]License; Usage
  • Context window 4,096 tokens tokens Primary source[2]Model Overview
  • Parameters 4BAs given in NVIDIA's model name; neither source states a separate parameter count. Primary source[1] [2]
  • Access API, On device, Open-weights downloadOffered as an NVIDIA NIM for cloud and on-device deployment (GeForce RTX PCs), hosted at build.nvidia.com, and later as weights on Hugging Face. Primary source[1]NVIDIA ACE introduces an SLM purpose-built for roleplaying [2]Model Overview; Usage [3]

Lineage

Predecessors

No known predecessor.

Successors

No known successor.

Based on

Variants and derived

None recorded.

Siblings

None recorded.

All ancestors

All descendants

None.

Variants

No variants recorded in this record.

Related AI Radar coverage

AI Radar coverage starts in June 2026; no coverage linked yet.

All model releases from NVIDIA on AI Radar →