Mistral NeMo

Available · Language

Mistral NeMo is a 12B-parameter language model that Mistral AI built with NVIDIA, released with base and instruct weights under Apache 2.0 and offered on la Plateforme and as an NVIDIA NIM. It has a 128k-token context window, introduced the Tekken tokenizer, and Mistral presents it as a drop-in replacement for Mistral 7B. [1] [2] [3] Primary source

Timeline of Mistral NeMo →

Claims and evidence

  • Released Primary source[1]page date; first paragraph [4]July 18, 2024 [5]release date [2]first paragraph
  • Deprecated Primary source[6]Deprecated & retired models table, row Mistral Nemo 12B [5]Deprecation date
  • Retired Primary source[6]Deprecated & retired models table, row Mistral Nemo 12B
  • Status AvailableApache 2.0 weights remain downloadable from the mistralai Hugging Face organisation; the API endpoint open-mistral-nemo-2407 was deprecated on 2026-05-22 and retired on 2026-07-31. Primary source[7] [3] [6]Deprecated & retired models table, row Mistral Nemo 12B
  • Replaced by Ministral 3Mistral names Ministral 3 8B, a variant of the Ministral 3 record. Primary source[6]Deprecated & retired models table, row Mistral Nemo 12B, Alternative [5]Replacement
  • Successor of Mistral 7B v0.3 EditorialMistral presents Mistral NeMo as a drop-in replacement for Mistral 7B and compares the two, without naming a Mistral 7B version; linking to v0.3, the Mistral 7B version current on la Plateforme at the time, is an editorial choice. Primary source[1]first paragraph; Instruction fine-tuning [3]Key features
  • Change · Size Has 12B parameters, up from 7B in Mistral 7B, while keeping a standard architecture that Mistral says lets it replace Mistral 7B as a drop-in.Compared with Mistral 7B v0.3 Primary source[1]
  • Change · Architecture Uses a new tokenizer, Tekken, based on Tiktoken, in place of the SentencePiece tokenizer of earlier Mistral models.Compared with Mistral 7B v0.3 Primary source[1]
  • Change · Efficiency Trained with quantisation awareness to allow FP8 inference, which Mistral says comes without a loss in performance.Compared with Mistral 7B v0.3 Primary source[1]
  • Change · Languages Designed for multilingual use; Mistral names eleven languages, among them Chinese, Japanese, Korean, Arabic and Hindi.Compared with Mistral 7B v0.3 Primary source[1]
  • Input text Primary source[5]Modalities
  • Output text Primary source[5]Modalities
  • Feature Function calling Primary source[1]Multilingual Model for the Masses [5]Features
  • Feature Multilingual Primary source[1]Multilingual Model for the Masses
  • Open weights YesBase and instruct checkpoints under Apache 2.0. Primary source[1]second paragraph; Links [7] [3]
  • Context window 128k tokens Primary source[1]first paragraph [5]Context
  • Parameters 12B Primary source[1]first paragraph [5]Weights table
  • Access API, Open-weights download Primary source[1]Links

Lineage

Predecessors

Successors

No known successor.

Based on

Not derived from another model.

Variants and derived

Siblings

None recorded.

All ancestors

All descendants

Variants

No variants recorded in this record.

Related AI Radar coverage

AI Radar coverage starts in June 2026; no coverage linked yet.

All model releases from Mistral AI on AI Radar →