Claims and evidence
- Successor of DeepSeek-V4-Flash-0731The change log calls V4 Flash and V4 Flash Vision Exp the previous-generation models, retired with this release, and routes their API names to V4.1-Flash; at that point deepseek-v4-flash served DeepSeek-V4-Flash-0731 (change log, 2026-07-31). Primary source[3]Date: 2026-09-10 [2]
- Change · Modality Processes images natively together with text, trained jointly from the start of pre-training; the previous Flash model accepted text only.Compared with DeepSeek-V4-Flash-0731 Primary source[1]Introduction
- Change · Architecture Switches to a Causal Encoder-Decoder design with 8B active parameters for input and 16B for output, plus CSA2 sparse attention and Engram conditional memory.Compared with DeepSeek-V4-Flash-0731 Primary source[1]Introduction
- Change · Training data Trained from scratch on a 45T-token multimodal corpus, rather than derived from a V4 checkpoint.Compared with DeepSeek-V4-Flash-0731 Primary source[1]
- Change · Efficiency According to DeepSeek, its KV cache needs about a quarter of the HBM and an eighth of the SSD storage of the previous generation.Compared with DeepSeek-V4-Flash-0731 Primary source[2]
- Input text, image Primary source[1]Introduction [4]Model Details: Features, Vision
- Output text Primary source[1]Introduction
- Feature Reasoning modeContinuously controllable reasoning effort (1 to 100) in the open model; thinking and non-thinking modes in the API. Primary source[1]Post-training [4]Model Details: Thinking mode
- Feature Function calling Primary source[4]Model Details: Features, Tool Calls
- Feature Long context Primary source[1]Introduction
- Open weights YesWeights under the MIT License. Primary source[1]License [2]
- Context window 1M tokens Primary source[1]Introduction [4]Model Details: Context length
- Parameters 552BMixture-of-Experts with 552B backbone parameters; 8B activated per token during prefill (input) and 16B during decode (output). The card also lists an Engram conditional memory of 196B parameters among additional components; whether it is counted in the 552B is not stated. Primary source[1]Introduction [2]
- Access API, Open-weights download Primary source[2] [1]License
Lineage
Predecessors
- DeepSeek-V4-Flash-0731 · 31 July 2026
Successors
No known successor.
Based on
Not derived from another model.
Variants and derived
None recorded.
Siblings
None recorded.
All ancestors
- DeepSeek LLM · 29 November 2023
- DeepSeek-V2 · 6 May 2024
- DeepSeek-V2-Chat-0628 · 28 June 2024
- DeepSeek-V2.5 · 5 September 2024
- DeepSeek-V2.5-1210 · 10 December 2024
- DeepSeek-V3 · 26 December 2024
- DeepSeek-V3-0324 · 24 March 2025
- DeepSeek-V3.1 · 21 August 2025
- DeepSeek-V3.1-Terminus · 22 September 2025
- DeepSeek-V3.2-Exp · 29 September 2025
- DeepSeek-V3.2 · 1 December 2025
- DeepSeek-V4-Flash · 24 April 2026
- DeepSeek-V4-Flash-0731 · 31 July 2026
All descendants
None.
Variants
No variants recorded in this record.
Related AI Radar coverage
- Lenovo's TianxiCode Agent and DeepSeek-V4.1-Flash Top SWE-bench-Live Lite at 71%, VerifiedDeepSeek · Pandaily ·
- Did Claude Haiku 5.5 Just DESTROY GPT-6.1 Luna & DeepSeek V4.1 Flash? (The 90% Price Cut Shockwave)DeepSeek · YouTube ·
- DeepSeek V4.1 Flash vs Haiku 5.5 vs Gemini FlashDeepSeek · https://tech-insider.org/ ·
- DeepSeek V4.1 Flash Q4 Requires 294 GB of RAM vs 152 GB for Q2DeepSeek · Geeky Gadgets ·
- US-China AI Gap Hits 3%, and DeepSeek V4.1 Flash Now Leads on Agentic Coding BenchmarksDeepSeek · Tech Times ·
- The 552B DeepSeek V4.1-Flash Model Offers A Peak Output Of 494 Tokens/Second When Powered By An At-Home Rig Spanning 4x NVIDIA DGX Spark UnitsDeepSeek · Wccftech ·
- DeepSeek V4.1 Flash: Two Rates for One Model, and a Legacy Id That Now Serves AnotherDeepSeek · Asian Movie Pulse ·