Claims and evidence
- Status AvailableWeights are still public on DeepSeek's Hugging Face organisation under the MIT License. In the DeepSeek API, deepseek-chat and deepseek-reasoner moved to DeepSeek-V3.1-Terminus on 2025-09-22. Primary source[2]License [3]Date: 2025-09-22
- Replaced by DeepSeek-V3.1-TerminusIn the DeepSeek API only: both API model names were upgraded to DeepSeek-V3.1-Terminus. The weights stay available. Primary source[3]Date: 2025-09-22
- Successor of DeepSeek-V3-0324deepseek-chat served DeepSeek-V3-0324 from 2025-03-24 and was upgraded to DeepSeek-V3.1 on 2025-08-21; the model card calls V3.1 an upgrade over the previous version and compares it with DeepSeek V3 0324. Primary source[3]Date: 2025-08-21; Date: 2025-03-24 [2]Introduction; Evaluation
- Derived from (continued pretraining) DeepSeek-V3DeepSeek-V3.1-Base was built on the original V3 base checkpoint through a two-phase long-context extension; DeepSeek-V3.1 was post-trained on top of it. Primary source[2]Introduction [1]Model Update
- Change · Reasoning One hybrid model serves both a thinking and a non-thinking mode; in the API these modes had been served by separate models, DeepSeek-V3-0324 and DeepSeek-R1-0528.Compared with DeepSeek-V3-0324 Primary source[3]Date: 2025-08-21 [2]Introduction
- Change · Tool use DeepSeek says post-training improved tool use and performance on agent tasks.Compared with DeepSeek-V3-0324 Primary source[2]Introduction
- Change · Efficiency DeepSeek says the thinking mode gives answers of comparable quality to DeepSeek-R1-0528 while responding more quickly.Compared with DeepSeek-V3-0324 Primary source[2]Introduction
- Change · Training data The base model received 840B tokens of continued pre-training on top of DeepSeek-V3 to extend its context length, in a 32K phase and a 128K phase.Compared with DeepSeek-V3 Primary source[1]Model Update [2]Introduction
- Feature Reasoning modeOne hybrid model with a thinking and a non-thinking mode; in the API deepseek-reasoner gave the thinking mode and deepseek-chat the non-thinking mode. Primary source[1] [2]Introduction
- Feature Function callingTool calls in non-thinking mode per the model card; strict function calling in the Beta API at launch. Primary source[1]API Update [2]ToolCall
- Open weights YesMIT License for the repository and the weights. Primary source[1]Model Update [2]License
- Context window 128K tokens Primary source[2]Model Downloads [1]API Update
- Parameters 671B total, 37B activatedMixture-of-experts model; the model card says the model structure is the same as DeepSeek-V3. Primary source[2]Model Downloads
- Access API, consumer app, Open-weights download Primary source[1]
Lineage
Predecessors
- DeepSeek-V3-0324 · 24 March 2025
Successors
No known successor.
Based on
- DeepSeek-V3 · 26 December 2024 · derived (continued pretraining)
Variants and derived
- DeepSeek-V3.1-Terminus · 22 September 2025 · revision
Siblings
None recorded.
All ancestors
- DeepSeek LLM · 29 November 2023
- DeepSeek-V2 · 6 May 2024
- DeepSeek-V2-Chat-0628 · 28 June 2024
- DeepSeek-V2.5 · 5 September 2024
- DeepSeek-V2.5-1210 · 10 December 2024
- DeepSeek-V3 · 26 December 2024
- DeepSeek-V3-0324 · 24 March 2025
All descendants
- DeepSeek-V3.1-Terminus · 22 September 2025
- DeepSeek-V3.2-Exp · 29 September 2025
- DeepSeek-V3.2 · 1 December 2025
- DeepSeek-V3.2-Speciale · 1 December 2025
- DeepSeek-V4-Flash · 24 April 2026
- DeepSeek-V4-Pro · 24 April 2026
- DeepSeek-V4-Flash-0731 · 31 July 2026
- DeepSeek-V4-Pro-0813 · 13 August 2026
- DeepSeek-V4-Flash-Vision-Exp · 21 August 2026
- DeepSeek-V4.1-Flash · 10 September 2026
Variants
DeepSeek-V3.1-Base
Same dates as DeepSeek-V3.1.
Related AI Radar coverage
AI Radar coverage starts in June 2026; no coverage linked yet.