DeepSeek-V3

Available · Language · Milestone

DeepSeek-V3 is a mixture-of-experts language model with 671B total and 37B activated parameters, offered in DeepSeek's web chat, through its API and as open weights together with the base model DeepSeek-V3-Base. DeepSeek pre-trained it on 14.8 trillion tokens and says it validated FP8 mixed-precision training at this scale for the first time. [1] [2]3. Model Downloads; 5. Chat Website & API Platform [3] Primary source

Timeline of DeepSeek-V3 →

Claims and evidence

  • Released Primary source[4]news sidebar: Introducing DeepSeek-V3 2024/12/26 [5]Date: 2024-12-26 [6]page date
  • Status AvailableWeights are still public on Hugging Face. In the DeepSeek API, deepseek-chat moved to DeepSeek-V3-0324 on 2025-03-24. Primary source[1] [7] [5]Date: 2025-03-24
  • Replaced by DeepSeek-V3-0324In the DeepSeek API only: deepseek-chat was upgraded to DeepSeek-V3-0324. The weights stay available. Primary source[5]Date: 2025-03-24
  • Successor of DeepSeek-V2.5-1210 EditorialEditorial link along the V line: DeepSeek-V3 is the next generation after the V2 series, whose last release was DeepSeek-V2.5-1210, and it replaced that model behind deepseek-chat. No DeepSeek page calls it the successor of V2.5-1210 in so many words. Primary source[4] [5]Date: 2024-12-26
  • Change · Size Total parameters rise to 671B with 37B activated per token, from the 236B total and 21B activated that DeepSeek lists for DeepSeek-V2 and V2.5.Compared with DeepSeek-V2.5-1210 Primary source[1]
  • Change · Architecture Adds an auxiliary-loss-free load-balancing strategy and a multi-token prediction training objective to the DeepSeek-V2 architecture.Compared with DeepSeek-V2.5-1210 Primary source[3]
  • Change · Efficiency DeepSeek says generation reaches 60 tokens per second, three times as fast as V2.Compared with DeepSeek-V2.5-1210 Primary source[4]
  • Change · Reasoning Post-training distils reasoning abilities from a DeepSeek-R1 series model into DeepSeek-V3.Compared with DeepSeek-V2.5-1210 Primary source[1]
  • Input textPresented as a language model; the announcement says multimodal support would come later. Primary source[1]1. Introduction [4]closing section
  • Output text Primary source[1]1. Introduction
  • Open weights YesCode under MIT; model use under the DeepSeek Model License, which supports commercial use. Primary source[2]3. Model Downloads; 7. License [1]7. License
  • Context window 128K tokens Primary source[2]3. Model Downloads [1]3. Model Downloads
  • Parameters 671B total, 37B activatedMixture-of-experts. The Hugging Face upload totals 685B including 14B of Multi-Token Prediction module weights. Primary source[2]1. Introduction; 3. Model Downloads [3]abstract [4]What's new in V3
  • Access consumer app, API, Open-weights download Primary source[2]3. Model Downloads; 5. Chat Website & API Platform [5]Date: 2024-12-26

Lineage

Predecessors

Successors

No known successor.

Based on

Not derived from another model.

Variants and derived

Siblings

None recorded.

All ancestors

All descendants

Variants

DeepSeek-V3-Base

Same dates as DeepSeek-V3.

  • Variant Inline variant in this record.Pre-trained base model, published on Hugging Face together with the chat model. The Hugging Face API gives its repository creation time as 2024-12-25, which no page states in text, so the variant has no dates of its own. Primary source[2]3. Model Downloads [7]

Related AI Radar coverage

AI Radar coverage starts in June 2026; no coverage linked yet.

All model releases from DeepSeek on AI Radar →