DeepSeek models
DeepSeek (Hangzhou DeepSeek Artificial Intelligence Co., Ltd.) is a Chinese AI company that says it was founded in 2023. It describes itself as a research team focused on fundamental model research with an open-source approach, and says it releases model weights under the MIT License with a technical report for each model. It offers its models through a web chat, mobile apps and an API platform.
Chronology
-
2026
6 models-
ReleasedDeepSeek Flash · Language, Multimodal, Reasoning DeepSeek-V4.1-Flash
DeepSeek-V4.1-Flash is a Mixture-of-Experts model with 552B backbone parameters that takes text and images and generates text, released on 2026-09-10 in the DeepSeek API as deepseek-flash and with…
- Processes images natively together with text, trained jointly from the start of pre-training; the previous Flash model accepted text only.
- Switches to a Causal Encoder-Decoder design with 8B active parameters for input and 16B for output, plus CSA2 sparse attention and Engram conditional memory.
Primary source Available · 7 news items -
ReleasedDeepSeek Flash · Language, Multimodal, Reasoning DeepSeek-V4-Flash-Vision-Exp
DeepSeek-V4-Flash-Vision-Exp is an experimental model that DeepSeek calls its first multimodal model in the DeepSeek-V4 family, built on the V4-Flash architecture with added visual modules. It…
- Adds image input next to text; images can be sent as base64, as external URLs or through the new Files API.
- Adds visual modules to the V4-Flash architecture, followed by continued training.
Primary source Available · no news linked -
ReleasedDeepSeek Pro · Language, Reasoning DeepSeek-V4-Pro-0813
DeepSeek-V4-Pro-0813 is the general-availability release of DeepSeek-V4-Pro, rolled out on 2026-08-13 in the DeepSeek app, web chat (Expert Mode) and API under the unchanged name deepseek-v4-pro…
- Adds an attached DSpark speculative decoding module on top of the preview's model structure.
- Natively supports the OpenAI Responses API format, adapted for Codex with a one-click setup script.
Primary source Available · no news linked -
ReleasedDeepSeek Flash · Language, Reasoning DeepSeek-V4-Flash-0731
DeepSeek-V4-Flash-0731 is the official, non-preview release of DeepSeek-V4-Flash, with open weights under the MIT License, served in the DeepSeek API from 2026-07-31 under the unchanged name…
- Ships with an attached DSpark speculative decoding module; DeepSeek says architecture and size are otherwise unchanged and only post-training was redone.
- Reasoning effort is set through a reasoning_effort parameter with three levels: low, high and max.
Primary source Available · no news linked -
ReleasedDeepSeek Flash · Language, Reasoning · Milestone DeepSeek-V4-Flash
DeepSeek-V4-Flash is an open-weights Mixture-of-Experts language model with 284B total and 13B active parameters, the smaller model of the DeepSeek-V4 preview released on 2026-04-24 alongside…
- Uses the V4 hybrid attention design (Compressed Sparse Attention plus Heavily Compressed Attention), Manifold-Constrained Hyper-Connections and the Muon optimizer.
- 284B total and 13B activated parameters, against 671B total and 37B activated for DeepSeek-V3.2-Base in the card's base-model table.
Primary source Available · no news linked -
ReleasedDeepSeek Pro · Language, Reasoning · Milestone DeepSeek-V4-Pro
DeepSeek-V4-Pro is an open-weights Mixture-of-Experts language model with 1.6T total and 49B active parameters, released on 2026-04-24 as the preview of the DeepSeek-V4 series alongside…
- Introduces hybrid attention (Compressed Sparse Attention plus Heavily Compressed Attention), Manifold-Constrained Hyper-Connections and training with the Muon optimizer.
- At a 1M-token context it needs about 27% of the single-token inference FLOPs and 10% of the KV cache of DeepSeek-V3.2, according to DeepSeek.
Primary source Available · no news linked
-
-
2025
9 models-
ReleasedDeepSeek-V · Language, Reasoning DeepSeek-V3.2
DeepSeek-V3.2 is a text-only language model in DeepSeek's V line with thinking and non-thinking modes, released on 1 December 2025 in the app, web chat and API, with open weights under the MIT…
- Integrates thinking directly into tool use, which DeepSeek says is a first for its models, and supports tool use in both thinking and non-thinking modes.
- Trained with synthesised agent data that DeepSeek says covers more than 1,800 environments and 85,000 complex instructions.
Primary source Available · no news linked -
ReleasedDeepSeek-V · Language, Reasoning DeepSeek-V3.2-Speciale
DeepSeek-V3.2-Speciale is a high-compute variant of DeepSeek-V3.2 designed for deep reasoning tasks, announced with it on 1 December 2025 and released with open weights under the MIT License. It…
- Trained only on reasoning data with a reduced length penalty in reinforcement learning, plus the DeepSeekMath-V2 dataset and reward method for mathematical proofs.
- Does not support tool calling.
Primary source Available · no news linked -
ReleasedDeepSeek-V · Language, Reasoning DeepSeek-V3.2-Exp
DeepSeek-V3.2-Exp is an experimental language model in DeepSeek's V line, released on 29 September 2025 in the app, web chat and API, with open weights under the MIT License. It adds DeepSeek Sparse…
- Adds DeepSeek Sparse Attention, a fine-grained sparse attention mechanism with a lightning indexer; DeepSeek says this is the only architectural change.
- DeepSeek says sparse attention makes long-context training and inference more efficient; it cut API prices by more than 50 percent with this release.
Primary source Available · no news linked -
ReleasedDeepSeek-V · Language, Reasoning DeepSeek-V3.1-Terminus
DeepSeek-V3.1-Terminus is an update of DeepSeek-V3.1, released on 22 September 2025 in the DeepSeek app, web chat and API, with open weights under the MIT License. DeepSeek says it keeps the…
- Less mixing of Chinese and English and fewer occasional abnormal characters in output, according to DeepSeek.
- DeepSeek reports better Code Agent and Search Agent performance; the search-agent template and tool set were updated.
Primary source Available · no news linked -
ReleasedDeepSeek-V · Language, Reasoning DeepSeek-V3.1
DeepSeek-V3.1 is a hybrid language model in DeepSeek's V line that offers a thinking and a non-thinking mode in one model, launched on 21 August 2025 in DeepSeek's web chat and API. Its weights and…
- One hybrid model serves both a thinking and a non-thinking mode; in the API these modes had been served by separate models, DeepSeek-V3-0324 and DeepSeek-R1-0528.
- DeepSeek says post-training improved tool use and performance on agent tasks.
Primary source Available · no news linked -
ReleasedDeepSeek-R · Language, Reasoning DeepSeek-R1-0528
DeepSeek-R1-0528 is a minor version upgrade of DeepSeek-R1, served behind deepseek-reasoner in the DeepSeek API from 28 May 2025, in the web chat and as open weights under the MIT License. DeepSeek…
- JSON output and function calling support are listed as part of the upgrade.
- Deeper reasoning, which DeepSeek attributes to more compute and algorithmic optimisation in post-training; complex tasks may use more tokens.
Primary source Available · no news linked -
ReleasedDeepSeek-V · Language DeepSeek-V3-0324
DeepSeek-V3-0324 is an updated version of DeepSeek-V3 with the same model structure, served behind deepseek-chat in the DeepSeek API from 24 March 2025, in the web chat and as open weights under the…
- Weights released under the MIT License instead of the DeepSeek Model License used for DeepSeek-V3.
- Improved reasoning, according to DeepSeek, with the model structure unchanged.
Primary source Available · no news linked -
ReleasedDeepSeek-R · Language, Reasoning · Milestone DeepSeek-R1
DeepSeek-R1 is a reasoning language model built on DeepSeek-V3-Base, released on 20 January 2025 in the DeepSeek web chat (DeepThink), in the API as deepseek-reasoner and as open weights under the…
- Offered in the API and as open weights, while DeepSeek-R1-Lite-Preview was available only in the DeepSeek web chat at launch.
- Code and weights under the MIT License, and DeepSeek allows the weights and API outputs to be used for fine-tuning and distillation.
Primary source Available · no news linked -
ReleasedDeepSeek-R · Language, Reasoning DeepSeek-R1-Zero
DeepSeek-R1-Zero is one of DeepSeek's first-generation reasoning models, trained with large-scale reinforcement learning directly on DeepSeek-V3-Base without supervised fine-tuning first, and…
- Reasoning is learned through large-scale reinforcement learning (GRPO) applied directly to DeepSeek-V3-Base, with no supervised fine-tuning stage first.
Primary source Available · no news linked
-
-
2024
7 models-
ReleasedDeepSeek-V · Language · Milestone DeepSeek-V3
DeepSeek-V3 is a mixture-of-experts language model with 671B total and 37B activated parameters, offered in DeepSeek's web chat, through its API and as open weights together with the base model…
- Total parameters rise to 671B with 37B activated per token, from the 236B total and 21B activated that DeepSeek lists for DeepSeek-V2 and V2.5.
- Adds an auxiliary-loss-free load-balancing strategy and a multi-token prediction training objective to the DeepSeek-V2 architecture.
Primary source Available · no news linked -
ReleasedDeepSeek-V · Language DeepSeek-V2.5-1210
DeepSeek-V2.5-1210 is an upgraded version of DeepSeek-V2.5, released on 10 December 2024 behind deepseek-chat in the DeepSeek API and as open weights on Hugging Face. DeepSeek presented it as the…
- DeepSeek reports improvements in maths, coding, writing and reasoning over DeepSeek-V2.5.
- DeepSeek says the update improves the user experience for file upload and web page summarisation.
Primary source Available · no news linked -
ReleasedDeepSeek-R · Language, Reasoning DeepSeek-R1-Lite-Preview
DeepSeek-R1-Lite-Preview is a preview of a DeepSeek reasoning model that went live in the DeepSeek web chat on 20 November 2024. DeepSeek highlighted that it shows its thought process in real time…
Primary source Status unknown · no news linked -
ReleasedDeepSeek-V · Language DeepSeek-V2.5
DeepSeek-V2.5 is a language model that DeepSeek made by merging DeepSeek-V2-0628 with its code model DeepSeek-Coder-V2-0724, to combine general conversation and coding in one model. It was released…
- Merges the general chat model DeepSeek-V2-0628 with the code model DeepSeek-Coder-V2-0724 into one model.
- One model now serves both the deepseek-chat and deepseek-coder API models, which DeepSeek kept for backward compatibility.
Primary source Available · no news linked -
ReleasedDeepSeek-V · Language DeepSeek-V2-Chat-0628
DeepSeek-V2-Chat-0628 is a revised DeepSeek-V2 chat model that went live in the DeepSeek API on 2024-06-28 as DeepSeek-V2-0628 and was later published as open weights under the DeepSeek Model…
- Built on the Coder-V2 base model instead of the DeepSeek-V2 base; DeepSeek says the switch was meant to improve code generation and reasoning.
- Instruction following in the system prompt improved, which the model card says helps tasks such as immersive translation and retrieval-augmented generation.
Primary source Available · no news linked -
ReleasedDeepSeek-V · Language · Milestone DeepSeek-V2
DeepSeek-V2 is a mixture-of-experts language model from DeepSeek with 236B total parameters, 21B of them activated per token, and a 128K-token context, released as open weights and offered on…
- Moves from a dense transformer to a mixture-of-experts design (DeepSeekMoE) and adds the new Multi-head Latent Attention.
- Context length of 128K tokens, up from the 4,096-token sequence length of DeepSeek LLM.
Primary source Available · no news linked -
ReleasedDeepSeek · Language DeepSeekMoE 16B
DeepSeekMoE 16B is an open-weights mixture-of-experts language model from DeepSeek with 16.4B total and about 2.8B activated parameters, released as a Base and a Chat checkpoint with the DeepSeekMoE…
- Uses a sparse mixture-of-experts architecture (DeepSeekMoE) instead of the dense transformer of DeepSeek LLM 7B.
- DeepSeek reports results comparable with its dense DeepSeek 7B, trained on the same corpus, with about 40% of the computation.
Primary source Available · no news linked
-
-
2023
1 model-
ReleasedDeepSeek · Language · Milestone DeepSeek LLM
DeepSeek LLM is a general-purpose language model series from DeepSeek, released as open weights in 7B and 67B sizes, each as a Base and a Chat model, under the DeepSeek Model License, which permits…
Primary source Available · no news linked
-
Lineage
-
DeepSeek LLM
- Successors
- DeepSeek-V2
-
DeepSeekMoE 16B
No recorded relations.
-
DeepSeek-V2
- Predecessors
- DeepSeek LLM
- Variants and derived
- DeepSeek-V2-Chat-0628 · revision
-
DeepSeek-V2-Chat-0628
- Based on
- DeepSeek-V2 · revision
- Variants and derived
- DeepSeek-V2.5 · derived: merge
- Siblings
- DeepSeek-V2 (base model) · inline variant of DeepSeek-V2; DeepSeek-V2-Chat · inline variant of DeepSeek-V2; DeepSeek-V2-Lite · inline variant of DeepSeek-V2; DeepSeek-V2-Lite-Chat · inline variant of DeepSeek-V2
-
DeepSeek-V2.5
- Based on
- DeepSeek-V2-Chat-0628 · derived: merge
- Variants and derived
- DeepSeek-V2.5-1210 · revision
-
DeepSeek-R1-Lite-Preview
- Successors
- DeepSeek-R1
-
DeepSeek-V2.5-1210
- Successors
- DeepSeek-V3
- Based on
- DeepSeek-V2.5 · revision
-
DeepSeek-V3
- Predecessors
- DeepSeek-V2.5-1210
- Variants and derived
- DeepSeek-R1 · derived: fine tune; DeepSeek-R1-Zero · derived: other; DeepSeek-V3-0324 · revision; DeepSeek-V3.1 · derived: continued pretraining
-
DeepSeek-R1
- Predecessors
- DeepSeek-R1-Lite-Preview
- Based on
- DeepSeek-V3 · derived: fine tune
- Variants and derived
- DeepSeek-R1-0528 · revision; MAI-DS-R1 (Microsoft) · derived: fine tune; Phi-4-mini-flash-reasoning (Microsoft) · derived: distillation; Phi-4-mini-reasoning (Microsoft) · derived: distillation; R1 1776 (Perplexity) · derived: fine tune; Sonar Reasoning Pro (Perplexity) · derived: other
-
DeepSeek-R1-Zero
- Based on
- DeepSeek-V3 · derived: other
-
DeepSeek-V3-0324
- Successors
- DeepSeek-V3.1
- Based on
- DeepSeek-V3 · revision
- Siblings
- DeepSeek-V3-Base · inline variant of DeepSeek-V3
-
DeepSeek-R1-0528
- Based on
- DeepSeek-R1 · revision
-
DeepSeek-V3.1
- Predecessors
- DeepSeek-V3-0324
- Based on
- DeepSeek-V3 · derived: continued pretraining
- Variants and derived
- DeepSeek-V3.1-Terminus · revision
-
DeepSeek-V3.1-Terminus
- Based on
- DeepSeek-V3.1 · revision
- Variants and derived
- DeepSeek-V3.2-Exp · derived: continued pretraining
- Siblings
- DeepSeek-V3.1-Base · inline variant of DeepSeek-V3.1
-
DeepSeek-V3.2-Exp
- Successors
- DeepSeek-V3.2
- Based on
- DeepSeek-V3.1-Terminus · derived: continued pretraining
-
DeepSeek-V3.2
- Predecessors
- DeepSeek-V3.2-Exp
- Successors
- DeepSeek-V4-Flash; DeepSeek-V4-Pro
- Variants and derived
- DeepSeek-V3.2-Speciale · variant
-
DeepSeek-V3.2-Speciale
- Based on
- DeepSeek-V3.2 · variant
-
DeepSeek-V4-Flash
- Predecessors
- DeepSeek-V3.2
- Variants and derived
- DeepSeek-V4-Flash-0731 · revision; DeepSeek-V4-Flash-Vision-Exp · derived: other
-
DeepSeek-V4-Pro
- Predecessors
- DeepSeek-V3.2
- Variants and derived
- DeepSeek-V4-Pro-0813 · revision
-
DeepSeek-V4-Flash-0731
- Successors
- DeepSeek-V4.1-Flash
- Based on
- DeepSeek-V4-Flash · revision
- Siblings
- DeepSeek-V4-Flash-Base · inline variant of DeepSeek-V4-Flash
-
DeepSeek-V4-Pro-0813
- Based on
- DeepSeek-V4-Pro · revision
- Siblings
- DeepSeek-V4-Pro-Base · inline variant of DeepSeek-V4-Pro
-
DeepSeek-V4-Flash-Vision-Exp
- Based on
- DeepSeek-V4-Flash · derived: other
-
DeepSeek-V4.1-Flash
- Predecessors
- DeepSeek-V4-Flash-0731
Families
-
DeepSeek
DeepSeek LLM · DeepSeekMoE 16B
-
DeepSeek-R
DeepSeek-R1-Lite-Preview · DeepSeek-R1 · DeepSeek-R1-Zero · DeepSeek-R1-0528
-
DeepSeek-V
DeepSeek-V2 · DeepSeek-V2-Chat-0628 · DeepSeek-V2.5 · DeepSeek-V2.5-1210 · DeepSeek-V3 · DeepSeek-V3-0324 · DeepSeek-V3.1 · DeepSeek-V3.1-Terminus · DeepSeek-V3.2-Exp · DeepSeek-V3.2 · DeepSeek-V3.2-Speciale
-
DeepSeek Flash
DeepSeek-V4-Flash · DeepSeek-V4-Flash-0731 · DeepSeek-V4-Flash-Vision-Exp · DeepSeek-V4.1-Flash
-
DeepSeek Pro
-
-
About the organisation
- Description DeepSeek (Hangzhou DeepSeek Artificial Intelligence Co., Ltd.) is a Chinese AI company that says it was founded in 2023. It describes itself as a research team focused on fundamental model research with an open-source approach, and says it releases model weights under the MIT License with a technical report for each model. It offers its models through a web chat, mobile apps and an API platform. Primary source[1] [2] [3] [4]
Company · https://www.deepseek.com/ · 23 models listed here
Sources
-
[1] deepseek-ai (DeepSeek)
Model hub page · DeepSeek
- Provenance
- Primary
- Availability
- Active
- Last checked
Original ↗ https://huggingface.co/deepseek-ai DeepSeek's verified organisation page on Hugging Face. Its organisation card (the README of https://huggingface.co/spaces/deepseek-ai/README) says DeepSeek (深度求索) was founded in 2023 and is a Chinese company dedicated to making AGI a reality. Living page; content as seen on 2026-10-01. -
[2] DeepSeek Privacy Policy
Other · DeepSeek
- Provenance
- Primary
- Availability
- Active
- Last checked
Original ↗ https://cdn.deepseek.com/policies/en-US/deepseek-privacy-policy.html Names Hangzhou DeepSeek Artificial Intelligence Co., Ltd., with its registered address in China, as the provider and controller of the services. The page shows Last Update: Feb 10, 2026. Living page. -
[3] Model Mechanism and Training Methods of DeepSeek
Documentation · DeepSeek
- Provenance
- Primary
- Availability
- Active
- Last checked
Original ↗ https://cdn.deepseek.com/policies/en-US/model-algorithm-disclosure.html Linked from the Transparency Center (https://www.deepseek.com/en/transparency/) as Model Principles and Training Methodology. The page states no date. It refers to the company as Hangzhou DeepSeek Artificial Intelligence Co., Ltd. -
[4] DeepSeek | Into the Unknown
Other · DeepSeek
- Provenance
- Primary
- Availability
- Active
- Last checked
Original ↗ https://www.deepseek.com/en/ English homepage, linked from the Chinese homepage; links DeepSeek Web, app downloads, DeepSeek Harness, the API platform and the API docs. On 2026-10-01 www.deepseek.com failed the TLS handshake from the checking machine, while https://deepseek.com/en/ served the same page; the site declares www.deepseek.com canonical. Living page. Checked on 2026-10-05 through deepseek.com without www (same page, HTTP 200); the www host failed the TLS handshake from the checking network.