DeepSeek-V4.1-Flash

Available · Language, Multimodal, Reasoning

DeepSeek-V4.1-Flash is a Mixture-of-Experts model with 552B backbone parameters that takes text and images and generates text, released on 2026-09-10 in the DeepSeek API as deepseek-flash and with open weights under the MIT License. DeepSeek calls it the smallest model of a new architecture family and says it was trained from scratch. [1]Introduction; License [2] Primary source

Timeline of DeepSeek-V4.1-Flash →

Claims and evidence

  • Released Primary source[2] [3]Date: 2026-09-10
  • Status Available Primary source[4]Model Details: Model version [1]License
  • Successor of DeepSeek-V4-Flash-0731The change log calls V4 Flash and V4 Flash Vision Exp the previous-generation models, retired with this release, and routes their API names to V4.1-Flash; at that point deepseek-v4-flash served DeepSeek-V4-Flash-0731 (change log, 2026-07-31). Primary source[3]Date: 2026-09-10 [2]
  • Change · Modality Processes images natively together with text, trained jointly from the start of pre-training; the previous Flash model accepted text only.Compared with DeepSeek-V4-Flash-0731 Primary source[1]Introduction
  • Change · Architecture Switches to a Causal Encoder-Decoder design with 8B active parameters for input and 16B for output, plus CSA2 sparse attention and Engram conditional memory.Compared with DeepSeek-V4-Flash-0731 Primary source[1]Introduction
  • Change · Training data Trained from scratch on a 45T-token multimodal corpus, rather than derived from a V4 checkpoint.Compared with DeepSeek-V4-Flash-0731 Primary source[1]
  • Change · Efficiency According to DeepSeek, its KV cache needs about a quarter of the HBM and an eighth of the SSD storage of the previous generation.Compared with DeepSeek-V4-Flash-0731 Primary source[2]
  • Input text, image Primary source[1]Introduction [4]Model Details: Features, Vision
  • Output text Primary source[1]Introduction
  • Feature Reasoning modeContinuously controllable reasoning effort (1 to 100) in the open model; thinking and non-thinking modes in the API. Primary source[1]Post-training [4]Model Details: Thinking mode
  • Feature Function calling Primary source[4]Model Details: Features, Tool Calls
  • Feature Long context Primary source[1]Introduction
  • Open weights YesWeights under the MIT License. Primary source[1]License [2]
  • Context window 1M tokens Primary source[1]Introduction [4]Model Details: Context length
  • Parameters 552BMixture-of-Experts with 552B backbone parameters; 8B activated per token during prefill (input) and 16B during decode (output). The card also lists an Engram conditional memory of 196B parameters among additional components; whether it is counted in the 552B is not stated. Primary source[1]Introduction [2]
  • Access API, Open-weights download Primary source[2] [1]License

Lineage

Predecessors

Successors

No known successor.

Based on

Not derived from another model.

Variants and derived

None recorded.

Siblings

None recorded.

All ancestors

All descendants

None.

Variants

No variants recorded in this record.

Related AI Radar coverage

All model releases from DeepSeek on AI Radar →