Phi-4-mini-flash-reasoning

Available · Language, Reasoning

Phi-4-mini-flash-reasoning is a 3.8-billion-parameter open model from Microsoft for math reasoning with a 64K-token context, released on 9 July 2025 on Azure AI Foundry, the NVIDIA API Catalog and Hugging Face. It uses the SambaY decoder-hybrid-decoder architecture, which combines Mamba state space layers, sliding window attention and Gated Memory Units, and targets edge and mobile use. [1] [2] Primary source

Timeline of Phi-4-mini-flash-reasoning →

Claims and evidence

  • Released Other sources give: [2]Training, Model: Release date The model card gives June 2025 as release date; the announcement of July 9, 2025 says the model is available from that day. Primary source[1]page date and first paragraph
  • Status AvailableWeights remain downloadable from Microsoft's Hugging Face organisation (checked 2026-10-01). The model is not listed in the Microsoft Foundry retirement schedule. Primary source[2]weights and licence (MIT)
  • Successor of Phi-4-mini-reasoningThe announcement refers to a predecessor that, like this model, is a 3.8B open model for math reasoning, and compares the model directly with Phi-4-mini-reasoning. It also says the model follows Phi-4-mini but uses a new hybrid architecture; the model card describes its own pre-training run of 5T tokens, so no derived-from link to Phi-4-mini is recorded. Primary source[1]Efficiency without compromise; Phi-4-mini-flash-reasoning benchmarks [3]Abstract: compared with Phi4-mini-Reasoning
  • Derived from (distillation) DeepSeek-R1 (DeepSeek) Primary source[2]
  • Change · Architecture Replaces the dense Transformer with the SambaY decoder-hybrid-decoder architecture, which uses Gated Memory Units.Compared with Phi-4-mini-reasoning Primary source[1]Efficiency without compromise
  • Change · Efficiency Microsoft says it decodes with higher throughput and lower latency than Phi-4-mini-reasoning, especially on long generations.Compared with Phi-4-mini-reasoning Primary source[1]
  • Change · Context length Context length of 64K tokens, down from 128K for Phi-4-mini-reasoning.Compared with Phi-4-mini-reasoning Primary source[2]Training, Model: Context length [4]Training, Model: Context length
  • Change · Training data Reasoning training ran without a reinforcement learning stage, according to Microsoft's paper.Compared with Phi-4-mini-reasoning Primary source[3]Abstract
  • Input text Primary source[2]Training, Model: Inputs
  • Output text Primary source[2]Training, Model: Outputs
  • Open weights Yes Primary source[2]licence: MIT [1]Efficiency without compromise
  • Context window 64K tokens Primary source[1]Efficiency without compromise [2]Training, Model: Context length
  • Parameters 3.8BThe announcement writes 3.8 billion parameters. Primary source[2]Model Quality; Training, Model: Architecture [1]Efficiency without compromise
  • Access API, cloud partner, Open-weights downloadAzure AI Foundry (api), NVIDIA API Catalog (cloud-partner) and Hugging Face (open-weights-download) at release. Primary source[1]first paragraph

Lineage

Predecessors

Successors

No known successor.

Based on

  • DeepSeek-R1 · DeepSeek · 20 January 2025 · derived (distillation)

Variants and derived

None recorded.

Siblings

None recorded.

All ancestors

All descendants

None.

Variants

No variants recorded in this record.

Related AI Radar coverage

AI Radar coverage starts in June 2026; no coverage linked yet.

All model releases from Microsoft on AI Radar →