DeepSeek-V4-Flash Specs Leaked: 284B Parameters, FP4 Mixed Precision, Native 1M Context

vllm_project · x · 2026-08-01

The vLLM Recipes page has unexpectedly leaked detailed technical documentation for the DeepSeek-V4-Flash model. As a new member of the V4 preview family, the model has a total parameter count of 284B with only 13B active parameters.

Core Architectural Highlights:

Deployment & Variants:

Pre-trained on 32T+ tokens, the model utilizes a two-stage post-training pipeline. Four checkpoint variants are provided, including a native FP4+FP8 mixed weights version and an NVFP4 version re-quantized by NVIDIA, optimized specifically for Blackwell GPUs.

Related event: DeepSeek-V4-Flash Architecture Leaked with Million-Token Context(3 posts)→

Original post →

More from Infra

Infra channel →