DeepSeek-V4-Pro launches: 1.6T-param MoE cuts inference FLOPs to 27% of V3.2
AccBalanced · x · 2026-08-14
DeepSeek officially released its flagship model DeepSeek-V4-Pro, a 1.6T-total / 49B-active MoE with hybrid CSA+HCA attention and manifold-constrained hyper-connections, achieving 27% of V3.2's per-token inference FLOPs and 10% of KV cache at 1M context. Pre-trained on 32T+ tokens with Muon optimizer, it supports FP4+FP8 mixed precision and native OpenAI Responses API, optimized for Codex. vLLM support is ready with no config rebuild. DeepSeek also open-sourced an agent harness compatible with any OpenAI-compatible endpoint.
Related event: DeepSeek-V4-Pro Released: 1.6T MoE, Lower Inference Cost(5 posts)→
More from Infra
- Orion-16B passes 100B tokens, largest LLM pretrained with decentralized compute — const_reborn · 2026-08-15
- Why no P2P for AI model downloads? Hugging Face single point of failure — deathcom65 · 2026-08-15
- Qwen3.8 gets Day-0 support from LightSeek, boosting inference performance by 30%+ — Alibaba_Qwen · 2026-08-15
- Is Self-Hosting LLMs on Preemptible Cloud Instances Cost-Effective? Reddit Discusses — Different-Monk5916 · 2026-08-15
- Musk: Orbital compute may be only way to scale AI by 2029 — elonmusk · 2026-08-15
- Qwen3.8-27B-FP8 on GH200: 10 concurrent streaming requests, first token in 10ms — MaziyarPanahi · 2026-08-15