DeepSeek-V4-Pro launches: 1.6T-param MoE cuts inference FLOPs to 27% of V3.2

AccBalanced · x · 2026-08-14

DeepSeek officially released its flagship model DeepSeek-V4-Pro, a 1.6T-total / 49B-active MoE with hybrid CSA+HCA attention and manifold-constrained hyper-connections, achieving 27% of V3.2's per-token inference FLOPs and 10% of KV cache at 1M context. Pre-trained on 32T+ tokens with Muon optimizer, it supports FP4+FP8 mixed precision and native OpenAI Responses API, optimized for Codex. vLLM support is ready with no config rebuild. DeepSeek also open-sourced an agent harness compatible with any OpenAI-compatible endpoint.

Related event: DeepSeek-V4-Pro Released: 1.6T MoE, Lower Inference Cost(5 posts)→

Original post →

More from Infra

Infra channel →