DeepSeek-V4-Pro

Available · Language, Reasoning · Milestone

DeepSeek-V4-Pro is an open-weights Mixture-of-Experts language model with 1.6T total and 49B active parameters, released on 2026-04-24 as the preview of the DeepSeek-V4 series alongside DeepSeek-V4-Flash. It was offered in the DeepSeek API, the web chat and on Hugging Face under the MIT License, with a 1M-token context and three reasoning modes. [1] [2]Introduction; Model Downloads; License [3] Primary source

Timeline of DeepSeek-V4-Pro →

Claims and evidence

  • Released Primary source[1] [4]Date: 2026-04-24 [3]General Information: Release date
  • Status AvailablePreview weights remain downloadable from the deepseek-ai Hugging Face organisation. In the DeepSeek API the model name deepseek-v4-pro has served the GA release DeepSeek-V4-Pro-0813 since 2026-08-13. Primary source[2]Model Downloads; License [4]Date: 2026-08-13
  • Replaced by DeepSeek-V4-Pro-0813In the DeepSeek API, app and web: the GA release took over the model name deepseek-v4-pro on 2026-08-13; its card says it supersedes the preview. Primary source[4]Date: 2026-08-13 [5]Introduction
  • Successor of DeepSeek-V3.2 EditorialEditorial link along the DeepSeek-V line: the V4 card and report compare efficiency and base-model results with DeepSeek-V3.2, and the report abstract speaks of unnamed predecessors, but no source names V4-Pro the successor of V3.2. Primary source[2]Introduction; Base Model evaluation [6]Abstract
  • Change · Architecture Introduces hybrid attention (Compressed Sparse Attention plus Heavily Compressed Attention), Manifold-Constrained Hyper-Connections and training with the Muon optimizer.Compared with DeepSeek-V3.2 Primary source[2]Introduction
  • Change · Efficiency At a 1M-token context it needs about 27% of the single-token inference FLOPs and 10% of the KV cache of DeepSeek-V3.2, according to DeepSeek.Compared with DeepSeek-V3.2 Primary source[2]Introduction
  • Change · Size 1.6T total and 49B activated parameters, against 671B total and 37B activated for DeepSeek-V3.2-Base in the card's base-model table.Compared with DeepSeek-V3.2 Primary source[2]Base Model evaluation
  • Change · Context length Supports a 1M-token context, which DeepSeek made the default across its official services with V4.Compared with DeepSeek-V3.2 Primary source[1]
  • Input text Primary source[3]Model properties: Modalities
  • Output text Primary source[3]Model properties: Modalities
  • Feature Reasoning modeNon-think, Think High and Think Max modes; the API offers thinking and non-thinking modes. Primary source[2]Instruct Model: reasoning effort modes [1]API is Available Today [3]Model properties: Modalities
  • Feature Long context Primary source[1] [2]Introduction
  • Open weights YesWeights and code under the MIT License. Primary source[2]License [3]Methods of distribution and licenses [1]
  • Context window 1M tokens Primary source[2]Model Downloads [1] [3]Model properties: Context length
  • Parameters 1.6T total / 49B activeMixture-of-Experts: 1.6T total parameters, 49B activated per token. Primary source[1] [2]Model Downloads [3]Model properties: Total model size
  • Access API, consumer app, Open-weights download Primary source[1] [3]Methods of distribution and licenses

Lineage

Predecessors

Successors

No known successor.

Based on

Not derived from another model.

Variants and derived

Siblings

None recorded.

All ancestors

All descendants

Variants

DeepSeek-V4-Pro-Base

Same dates as DeepSeek-V4-Pro.

  • Variant Inline variant in this record.Pre-trained base checkpoint (FP8 mixed precision), listed with the preview release at https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-Base. Primary source[2]Model Downloads

Related AI Radar coverage

AI Radar coverage starts in June 2026; no coverage linked yet.

All model releases from DeepSeek on AI Radar →