
DeepSeek
AI · Infrastructure
DeepSeek Unveils V4 Preview: 1M Token Context MoE Models
DeepSeek launched preview versions of DeepSeek-V4-Pro and V4-Flash, both Mixture-of-Experts models with 1M token context windows, open-source weights, and API access. They lead open models in agentic coding, reasoning, and efficiency on Huawei chips.
DeepSeek released preview versions of its V4 model series on April 24, 2026, featuring DeepSeek-V4-Pro with 1.6 trillion total parameters and 49 billion active, and DeepSeek-V4-Flash with 284 billion total parameters and 13 billion active. Both support a standard 1 million token context window, trained on 33 trillion and 32 trillion tokens respectively, using innovations like token-wise compression, DeepSeek Sparse Attention, and hybrid attention mechanisms. These reduce inference FLOPs to 27% and KV cache to 10% of predecessor V3.2 levels for Pro at 1M context, enabling cost-effective long-context processing. The models offer thinking and non-thinking modes, with Pro-Max for high-reasoning tasks, and are available via API compatible with OpenAI and Anthropic formats, plus open-source weights on Hugging Face.