DeepSeek-V4-Flash

Available · Language, Reasoning · Milestone

DeepSeek-V4-Flash is an open-weights Mixture-of-Experts language model with 284B total and 13B active parameters, the smaller model of the DeepSeek-V4 preview released on 2026-04-24 alongside DeepSeek-V4-Pro. DeepSeek presents it as the faster, more economical option; it was offered in the API, the web chat and under the MIT License, with a 1M-token context. [1] [2]Introduction; Model Downloads; License Primary source

Timeline of DeepSeek-V4-Flash →

Claims and evidence

  • Released Primary source[1] [3]Date: 2026-04-24 [4]General Information: Release date
  • Status AvailablePreview weights remain downloadable from the deepseek-ai Hugging Face organisation. In the DeepSeek API the preview was replaced by DeepSeek-V4-Flash-0731 under the unchanged model name deepseek-v4-flash on 2026-07-31. Primary source[2]Model Downloads; License [3]Date: 2026-07-31
  • Replaced by DeepSeek-V4-Flash-0731In the DeepSeek API: DeepSeek-V4-Flash-0731 took over the model name deepseek-v4-flash on 2026-07-31; its card says it supersedes the preview. Primary source[3]Date: 2026-07-31 [5]Introduction
  • Successor of DeepSeek-V3.2 EditorialEditorial link: from 2026-04-24 the legacy API names deepseek-chat and deepseek-reasoner, which had served DeepSeek-V3.2 since 2025-12-01, pointed to the non-thinking and thinking modes of deepseek-v4-flash; the card also compares base results with DeepSeek-V3.2. No source calls V4-Flash the successor of V3.2. Primary source[3]Date: 2026-04-24; Date: 2025-12-01 [2]Base Model evaluation
  • Change · Architecture Uses the V4 hybrid attention design (Compressed Sparse Attention plus Heavily Compressed Attention), Manifold-Constrained Hyper-Connections and the Muon optimizer.Compared with DeepSeek-V3.2 Primary source[2]Introduction
  • Change · Size 284B total and 13B activated parameters, against 671B total and 37B activated for DeepSeek-V3.2-Base in the card's base-model table.Compared with DeepSeek-V3.2 Primary source[2]Base Model evaluation
  • Change · Context length Supports a 1M-token context, which DeepSeek made the default across its official services with V4.Compared with DeepSeek-V3.2 Primary source[1]
  • Change · Availability Took over the legacy API names deepseek-chat and deepseek-reasoner from DeepSeek-V3.2, with those names scheduled to be discontinued on 2026-07-24.Compared with DeepSeek-V3.2 Primary source[3]Date: 2026-04-24; Date: 2025-12-01
  • Input text Primary source[4]Model properties: Modalities
  • Output text Primary source[4]Model properties: Modalities
  • Feature Reasoning modeNon-think, Think High and Think Max modes; the API offers thinking and non-thinking modes. Primary source[2]Instruct Model: reasoning effort modes [1]API is Available Today [4]Model properties: Modalities
  • Feature Long context Primary source[1] [2]Introduction
  • Open weights YesWeights and code under the MIT License. Primary source[2]License [4]Methods of distribution and licenses [1]
  • Context window 1M tokens Primary source[2]Model Downloads [1] [4]Model properties: Context length
  • Parameters 284B total / 13B activeMixture-of-Experts: 284B total parameters, 13B activated per token. The V4 model card PDF on the Transparency Center gives 285B total parameters instead. Primary source[1] [2]Model Downloads [6]Abstract
  • Access API, consumer app, Open-weights download Primary source[1] [4]Methods of distribution and licenses

Lineage

Predecessors

Successors

No known successor.

Based on

Not derived from another model.

Variants and derived

Siblings

None recorded.

All ancestors

All descendants

Variants

DeepSeek-V4-Flash-Base

Same dates as DeepSeek-V4-Flash.

  • Variant Inline variant in this record.Pre-trained base checkpoint (FP8 mixed precision), listed with the preview release at https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-Base. Primary source[2]Model Downloads

Related AI Radar coverage

AI Radar coverage starts in June 2026; no coverage linked yet.

All model releases from DeepSeek on AI Radar →