DeepSeek-V4-Flash Specs Leaked: 284B Parameters, FP4 Mixed Precision, Native 1M Context
vllm_project · x · 2026-08-01
The vLLM Recipes page has unexpectedly leaked detailed technical documentation for the DeepSeek-V4-Flash model. As a new member of the V4 preview family, the model has a total parameter count of 284B with only 13B active parameters.
Core Architectural Highlights:
- Hybrid Attention Stack: Combines Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA).
- Extreme Memory Optimization: Paired with Manifold-Constrained Hyper-Connections, it requires only 27% of V3.2's per-token inference FLOPs and 10% of its KV Cache at a 1 million token context length.
- Mixed Precision Storage: MoE expert weights are stored in FP4, while remaining parameters (attention, router, etc.) use FP8.
- Inference Acceleration: Supports Multi-Token Prediction (MTP) and DSpark speculative decoding.
Deployment & Variants:
Pre-trained on 32T+ tokens, the model utilizes a two-stage post-training pipeline. Four checkpoint variants are provided, including a native FP4+FP8 mixed weights version and an NVFP4 version re-quantized by NVIDIA, optimized specifically for Blackwell GPUs.
Related event: DeepSeek-V4-Flash Architecture Leaked with Million-Token Context(3 posts)→
More from Infra
- Paper Share: How Chunked Prefill Improves LLM Serving Efficiency — Abhishekcur · 2026-08-01
- The Compute Bottlenecks of Agentic AI: Inference vs. Execution — charles_irl · 2026-08-01
- AI Market Correction Warning: Extreme Leverage in Memory Chips and Record Investor Debt — binarybits · 2026-08-01
- DeepSeek-V4-Flash Quantized on A100: Uses Only 15.8GB VRAM at 16 tok/s — Different-Pickle1021 · 2026-08-01
- Micron and Hynix Cautious on Capacity Due to Memory Cycle Scars, Price Hikes Signal Expansion — _sholtodouglas · 2026-08-01
- Running Kimi K3 on a B300: 450 tokens/s for $46k/month — casper_hansen_ · 2026-08-01