DeepSeek-V2

Available · Language · Milestone

DeepSeek-V2 is a mixture-of-experts language model from DeepSeek with 236B total parameters, 21B of them activated per token, and a 128K-token context, released as open weights and offered on DeepSeek's chat website and API. It introduced Multi-head Latent Attention, which compresses the key-value cache; smaller DeepSeek-V2-Lite models followed on 2024-05-16. [1]2. News; 3. Model Downloads; 6. Chat Website; 7. API Platform [2]abstract [3]2. News Primary source

Timeline of DeepSeek-V2 →

Claims and evidence

  • Released The technical report was posted on arXiv the next day, 2024-05-07. Primary source[1]2. News: 2024.05.06 [3]2. News
  • deepseek-chat upgraded to DeepSeek-V2-0517 Primary source[4]Date: 2024-05-17
  • Status AvailableWeights are still public and ungated on Hugging Face. In the DeepSeek API, deepseek-chat moved to DeepSeek-V2-0517 on 2024-05-17 and to DeepSeek-V2-0628 on 2024-06-28. Primary source[5] [6] [4]Date: 2024-05-17; Date: 2024-06-28
  • Replaced by DeepSeek-V2-Chat-0628In the DeepSeek API only: deepseek-chat was upgraded to DeepSeek-V2-0628. The weights stay available. Primary source[4]Date: 2024-06-28
  • Successor of DeepSeek LLMThe report calls DeepSeek 67B DeepSeek's previous release and reuses its tokenizer and data processing; the README labels it DeepSeek-V1. Primary source[2]section 1, Introduction; section 3.1.1, Data Construction; section 3.2.2 [1]4. Evaluation Results, DeepSeek-V1 (Dense-67B)
  • Change · Architecture Moves from a dense transformer to a mixture-of-experts design (DeepSeekMoE) and adds the new Multi-head Latent Attention.Compared with DeepSeek LLM Primary source[2]
  • Change · Context length Context length of 128K tokens, up from the 4,096-token sequence length of DeepSeek LLM.Compared with DeepSeek LLM Primary source[1]3. Model Downloads [7]2. Model Downloads, Sequence Length
  • Change · Training data Pre-trained on 8.1 trillion tokens; DeepSeek says the corpus holds more data than that of DeepSeek 67B, especially Chinese data, at higher quality.Compared with DeepSeek LLM Primary source[2]section 3.1.1, Data Construction
  • Change · Efficiency DeepSeek reports lower training cost, a much smaller key-value cache and higher maximum generation throughput than DeepSeek 67B.Compared with DeepSeek LLM Primary source[1]1. Introduction [2]abstract
  • Open weights YesDeepSeek Model License; commercial use permitted. Primary source[1]3. Model Downloads; 9. License [5]
  • Context window 128k tokens Primary source[1]3. Model Downloads [2]abstract
  • Parameters 236B total, 21B activated Primary source[1]1. Introduction; 3. Model Downloads [2]abstract
  • Access Open-weights download, consumer app, API Primary source[1]3. Model Downloads; 6. Chat Website; 7. API Platform

Lineage

Predecessors

Successors

No known successor.

Based on

Not derived from another model.

Variants and derived

Siblings

None recorded.

All ancestors

All descendants

Variants

DeepSeek-V2 (base model)

Same dates as DeepSeek-V2.

  • Variant Inline variant in this record.Published on Hugging Face as deepseek-ai/DeepSeek-V2; DeepSeek uses the plain name DeepSeek-V2 for it. Primary source[1]3. Model Downloads [5]

DeepSeek-V2-Chat

Same dates as DeepSeek-V2.

  • Variant Inline variant in this record.The released chat model is the RL version (DeepSeek-V2-Chat (RL)). Primary source[1]3. Model Downloads [6]

Also known as: DeepSeek-V2 Chat

DeepSeek-V2-Lite

  • Released Primary source[1]2. News: 2024.05.16 [3]2. News
  • Variant Inline variant in this record.Trained from scratch on 5.7T tokens, per the model card. Primary source[3]

Also known as: DeepSeek V2 Lite

Differs in:

  • Context window: 32k tokens Evidence not assessed [1]3. Model Downloads
  • Parameters: 16B total, 2.4B activated Evidence not assessed [1]3. Model Downloads [3]1. Introduction

DeepSeek-V2-Lite-Chat

  • Released Primary source[1]2. News: 2024.05.16 [8]2. News
  • Variant Inline variant in this record.Chat model made with SFT only (DeepSeek-V2-Lite-Chat (SFT)). Primary source[8]

Differs in:

  • Context window: 32k tokens Evidence not assessed [1]3. Model Downloads
  • Parameters: 16B total, 2.4B activated Evidence not assessed [1]3. Model Downloads

Related AI Radar coverage

AI Radar coverage starts in June 2026; no coverage linked yet.

All model releases from DeepSeek on AI Radar →