Mistral Small 4

Available · Language, Multimodal, Reasoning · Milestone

Mistral Small 4 is a mixture-of-experts model from Mistral AI with 119B total parameters, released on 16 March 2026 as Apache 2.0 open weights and on Mistral's API. Mistral describes it as its first model to combine the reasoning, multimodal and agentic coding capabilities of Magistral, Pixtral and Devstral, with reasoning effort set per request. [1] [2] [3] Primary source

Timeline of Mistral Small 4 →

Claims and evidence

  • Released Primary source[1] [4]March 16 (2026) [3]page header date
  • Status Available Primary source[3]GA badge [5]Generalist models [2]
  • Successor of Mistral Small 3.2 EditorialEditorial link along the Mistral Small line: the announcement calls Small 4 the next major release in the Mistral Small family and compares its non-reasoning mode with the chat style of Mistral Small 3.2, the most recent Small release before it, without naming a predecessor; the models overview names Small 4 as the alternative for the retired Small 3.2. Primary source[1]first paragraph; Reasoning on demand [5]Deprecated & retired models table, Alternative column
  • Change · Reasoning Adds reasoning effort set per request (none or high) in the same model, instead of a separate Magistral reasoning model.Compared with Mistral Small 3.2 Primary source[1]Reasoning on demand
  • Change · Architecture Moves to a mixture-of-experts design with 128 experts, 4 active per token, and 119B total parameters.Compared with Mistral Small 3.2 Primary source[2]Key Features
  • Change · Context length Has a context window of 256k tokens.Compared with Mistral Small 3.2 Primary source[3]Context
  • Change · Efficiency Mistral reports 40 percent less end-to-end completion time in a latency-optimized setup and three times more requests per second in a throughput-optimized setup than Small 3.Compared with Mistral Small 3 Primary source[1]
  • Input text, image Primary source[3]Modalities [1]Key architectural details
  • Output text Primary source[3]Modalities [2]Key Features
  • Feature Reasoning modeConfigurable per request with the reasoning_effort parameter (none or high). Primary source[1]Reasoning on demand [2]Recommended Settings
  • Feature Function calling Primary source[3]Features: Function Calling
  • Open weights YesApache 2.0. Primary source[2] [1]
  • Context window 256k tokens Primary source[3]Context [1]Key architectural details
  • Parameters 119B total, 6.5B activeMixture of experts with 128 experts, 4 active per token. The announcement gives 6B active parameters per token (8B including embedding and output layers); the docs and the model card give 6.5B active. Primary source[3]description [2]Key Features
  • Access API, Open-weights download, cloud partnerMistral API and AI Studio, Hugging Face, and NVIDIA (build.nvidia.com and NVIDIA NIM) at launch. Primary source[1]Availability

Lineage

Predecessors

Successors

No known successor.

Based on

Not derived from another model.

Variants and derived

None recorded.

Siblings

None recorded.

All ancestors

All descendants

None.

Variants

No variants recorded in this record.

Related AI Radar coverage

All model releases from Mistral AI on AI Radar →