Claims and evidence
- Status AvailablePreview weights remain downloadable from the deepseek-ai Hugging Face organisation. In the DeepSeek API the model name deepseek-v4-pro has served the GA release DeepSeek-V4-Pro-0813 since 2026-08-13. Primary source[2]Model Downloads; License [4]Date: 2026-08-13
- Replaced by DeepSeek-V4-Pro-0813In the DeepSeek API, app and web: the GA release took over the model name deepseek-v4-pro on 2026-08-13; its card says it supersedes the preview. Primary source[4]Date: 2026-08-13 [5]Introduction
- Successor of DeepSeek-V3.2 EditorialEditorial link along the DeepSeek-V line: the V4 card and report compare efficiency and base-model results with DeepSeek-V3.2, and the report abstract speaks of unnamed predecessors, but no source names V4-Pro the successor of V3.2. Primary source[2]Introduction; Base Model evaluation [6]Abstract
- Change · Architecture Introduces hybrid attention (Compressed Sparse Attention plus Heavily Compressed Attention), Manifold-Constrained Hyper-Connections and training with the Muon optimizer.Compared with DeepSeek-V3.2 Primary source[2]Introduction
- Change · Efficiency At a 1M-token context it needs about 27% of the single-token inference FLOPs and 10% of the KV cache of DeepSeek-V3.2, according to DeepSeek.Compared with DeepSeek-V3.2 Primary source[2]Introduction
- Change · Size 1.6T total and 49B activated parameters, against 671B total and 37B activated for DeepSeek-V3.2-Base in the card's base-model table.Compared with DeepSeek-V3.2 Primary source[2]Base Model evaluation
- Change · Context length Supports a 1M-token context, which DeepSeek made the default across its official services with V4.Compared with DeepSeek-V3.2 Primary source[1]
- Input text Primary source[3]Model properties: Modalities
- Output text Primary source[3]Model properties: Modalities
- Feature Reasoning modeNon-think, Think High and Think Max modes; the API offers thinking and non-thinking modes. Primary source[2]Instruct Model: reasoning effort modes [1]API is Available Today [3]Model properties: Modalities
- Feature Long context Primary source[1] [2]Introduction
- Open weights YesWeights and code under the MIT License. Primary source[2]License [3]Methods of distribution and licenses [1]
- Context window 1M tokens Primary source[2]Model Downloads [1] [3]Model properties: Context length
- Parameters 1.6T total / 49B activeMixture-of-Experts: 1.6T total parameters, 49B activated per token. Primary source[1] [2]Model Downloads [3]Model properties: Total model size
- Access API, consumer app, Open-weights download Primary source[1] [3]Methods of distribution and licenses
Lineage
Predecessors
- DeepSeek-V3.2 · 1 December 2025
Successors
No known successor.
Based on
Not derived from another model.
Variants and derived
- DeepSeek-V4-Pro-0813 · 13 August 2026 · revision
Siblings
None recorded.
All ancestors
- DeepSeek LLM · 29 November 2023
- DeepSeek-V2 · 6 May 2024
- DeepSeek-V2-Chat-0628 · 28 June 2024
- DeepSeek-V2.5 · 5 September 2024
- DeepSeek-V2.5-1210 · 10 December 2024
- DeepSeek-V3 · 26 December 2024
- DeepSeek-V3-0324 · 24 March 2025
- DeepSeek-V3.1 · 21 August 2025
- DeepSeek-V3.1-Terminus · 22 September 2025
- DeepSeek-V3.2-Exp · 29 September 2025
- DeepSeek-V3.2 · 1 December 2025
All descendants
- DeepSeek-V4-Pro-0813 · 13 August 2026
Variants
DeepSeek-V4-Pro-Base
Same dates as DeepSeek-V4-Pro.
- Variant Inline variant in this record.Pre-trained base checkpoint (FP8 mixed precision), listed with the preview release at https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-Base. Primary source[2]Model Downloads
Related AI Radar coverage
AI Radar coverage starts in June 2026; no coverage linked yet.