DeepSeek-V4 preview: million-token context at 27% of V3 FLOPs and 10% of KV cache

DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence

DeepSeek-AI, Anyi Xu, Bangcai Lin, Bing Xue, Bingxuan Wang, Bingzheng Xu, Bochao Wu, Bowei Zhang, Chaofan Lin, Chen Dong, Chenchen Ling, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chengyu Hou, Chenhao Xu, Chenze Shao, Chong Ruan, Conner Sun, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Donghao Li, Dongjie Ji, Erhang Li, Fang Wei, Fangyun Lin, Fangzhou Yuan, Feiyu Xia, Fucong Dai, Guangbo Hao, Guanting Chen, Guoai Cao, Guolai Meng, Guowei Li, Han Yu, Han Zhang, Hanwei Xu, Hao Li, Haofen Liang, Haoling Zhang, Haoming Luo, Haoran Wei, Haotian Yuan, Haowei Zhang, Haowen Luo, Haoyu Chen, Haozhe Ji, Hengqing Zhang, Honghui Ding, Hongxuan Tang, Huanqi Cao, Huazuo Gao, Hui Qu, Hui Zeng, J Yang, JQ Zhu, Jia Luo, Jia Song, Jia Yu, Jialiang Huang, Jialu Cai, Jian Liang, Jiangting Zhou, Jiasheng Ye, Jiashi Li, Jiaxin Xu, Jiewen Hu, Jieyu Yang, Jin Chen, Jin Yan, Jingchang Chen, Jingli Zhou, Jingting Xiang, Jingyang Yuan, Jingyuan Cheng, Jingzi Zhou, Jinhua Zhu, Jiping Yu, Joseph Sun, Jun Ran, Junguang Jiang, Junjie Qiu, Junlong Li, Junmin Zheng, Junxiao Song, Kai Dong, Kaige Gao, Kang Guan, Kexing Zhou, Kezhao Huang, Kuai Yu, Lean Wang, Lecong Zhang, Lei Wang, Leyi Xia, Li Zhang, Liang Zhao, Lihua Guo, Lingxiao Luo, Linwang Ma, Linyan Zhu, Litong Wang, Liyu Cai, Liyue Zhang, Longhao Chen, MS Di, MY Xu, Max Mei, Miaojun Wang, Mingchuan Zhang, Minghua Zhang, Minghui Tang, Mingming Li, Mingxu Zhou, Minmin Han, Ning Wang, Panpan Huang, Panpan Wang, Peixin Cong, Peiyi Wang, Peng Zhang, Qiancheng Wang, Qihao Zhu, Qingyang Li, Qinyu Chen, Qiushi Du, Qiwei Jiang, Rui Tian, Ruifan Xu, Ruijie Lu, Ruiling Xu, Ruiqi Ge, Ruisong Zhang, Ruizhe Pan, Runji Wang, Runqian Chen, Runqiu Yin, Runxin Xu, Ruomeng Shen, Ruoyu Zhang, Ruyi Chen, SH Liu, Shanghao Lu, Shangmian Sun, Shangyan Zhou, Shanhuang Chen, Shaofei Cai, Shaoheng Nie, Shaoqing Wu, Shaoyuan Chen, Shengding Hu, Shengyu Liu, Shiqiang Hu, Shirong Ma, Shiyu Wang, Shuiping Yu, Shunfeng Zhou, Shuting Pan, Shuying Yu, Songyang Zhou, Tao Ni, Tao Yun, Tian Jin, Tian Pei, Tian Ye, Tianle Lin, Tianran Ji, Tianyi Cui, Tianyuan Yue, Tingting Yu, Tun Wang, W Zhang, WL Xiao, Wangding Zeng, Wei An, Weilin Zhao, Wen Liu, Wenfeng Liang, Wenjie Pang, Wenjing Luo, Wenjing Yao, Wenjun Gao, Wenkai Yang, Wenlve Huang, Wenqing Hou, Wentao Zhang, Wenting Ma, Xi Gao, Xiang He, Xiangwen Wang, Xianzu Wang, Xiao Bi, Xiaodong Liu, Xiaohan Wang, Xiaokang Chen, Xiaokang Zhang, Xiaotao Nie, Xiaowen Sun, Xiaoxiang Wang, Xin Cheng, Xin Liu, Xin Xie, Xingchao Liu, Xingchen Liu, Xingkai Yu, Xingyou Li, Xinyu Yang, Xinyu Zhang, Xu Chen, Xuanyu Wang, Xuecheng Su, Xueyin Chen, Xuheng Lin, Xuwei Fu, YC Yan, YQ Wang, YW Ma, Yanfeng Luo, Yang Zhang, Yanhong Xu, Yanru Ma, Yanwen Huang, Yao Li, Yao Li, Yao Xu, Yao Zhao, Yaofeng Sun, Yaohui Wang, Yi Qian, Yi Shao, Yi Yu, Yichao Zhang, Yifan Ding, Yifan Shi, Yijia Wu, Yiliang Xiong, Yiling Ma, Ying He, Ying Tang, Ying Zhou, Yingjia Luo, Yinmin Zhong, Yishi Piao, Yisong Wang, Yixiang Zhang, Yixiao Chen, Yixuan Tan, Yixuan Wei, Yiyang Ma, Yiyuan Liu, Yonglun Yang, Yongqiang Guo, Yongtong Wu, Yu Wu, YuKun Li, Yuan Cheng, Yuan Ou, Yuanfan Xu, Yuanhao Li, Yuduan Wang, Yuehan Yang, Yuer Xu, Yuhan Wu, Yuhao Meng, Yuheng Zou, Yukun Zha, Yunfan Xiong, Yupeng Chen, Yuping Lin, Yuqian Cao, Yuqian Wang, Yushun Zhang, Yuting Yan, Yutong Lin, Yuxian Gu, Yuxiang Luo, Yuxiang You, Yuxuan Liu, Yuxuan Zhou, Yuyang Zhou, Yuzhen Huang, ZF Wu, Zehao Wang, Zehua Zhao, Zehui Ren, Zekai Zhang, Zhangli Sha, Zhe Fu, Zhe Ju, Zhean Xu, Zhenda Xie, Zhengyan Zhang, Zheren Gao, Zhewen Hao, Zhibin Gou, Zhicheng Ma, Zhigang Yan, Zhihong Shao, Zhixian Huang, Zhixuan Chen, Zhiyu Wu, Zhizhou Ren, Zhongyu Wu, Zhuoshu Li, Zhuping Zhang, Zian Xu, Zihao Wang, Zihua Qu, Zihui Gu, Zijia Zhu, Zilin Li, Zipeng Zhang, Ziwei Xie, Ziyi Gao, Ziyi Wan, Zizheng Pan, Zongqing Yao

cs.CL, cs.AI

2026-04-26

DeepSeek releases the V4 preview (Pro 1.6T/49B, Flash 284B/13B, both natively 1M-context). With hybrid CSA+HCA sparse attention, manifold-constrained hyper-connections, and the Muon optimizer, it pushes per-token FLOPs at million-token context to 27% of V3.2 and KV cache to 10%. V4-Pro-Max sets new open-source marks on knowledge and reasoning.

What problem this solves

Test-time scaling from reasoning models stretches context ever longer, while vanilla attention's compute grows quadratically with sequence length, making million-token contexts prohibitive. DeepSeek's V4 preview targets exactly this: break the ultra-long-context efficiency barrier without giving up capability, and make million-token context routine.

Method

V4 keeps DeepSeekMoE and multi-token prediction from V3 and changes three things.

A hybrid attention interleaves Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA): CSA compresses every m tokens of KV cache into one entry and then has each query attend to only k entries via a Lightning Indexer top-k selection, while HCA compresses more aggressively (m' far larger than m) but keeps dense attention; both add a small sliding window for local detail. Manifold-Constrained Hyper-Connections (mHC) upgrade residual connections by constraining the residual map to the Birkhoff polytope of doubly stochastic matrices, bounding the spectral norm at 1 so signal propagation is non-expansive, which fixes the numerical instability of stacked hyper-connections; projection is by Sinkhorn-Knopp. The Muon optimizer speeds convergence and stabilizes training.

Training trillion-parameter MoE is unstable; two tricks hold it together. Anticipatory Routing decouples the routing network from the backbone (routing uses historical parameters, precomputed and cached, and is applied dynamically only on loss spikes, at negligible overall cost). SwiGLU Clamping clamps the SwiGLU linear term to [-10, 10] and caps the gate at 10 to kill outliers. Post-training replaces the mixed-RL stage entirely with on-policy distillation to merge domain experts, and uses a Generative Reward Model (the actor judges itself) instead of a scalar reward model.

Results

V4-Pro (1.6T total, 49B activated) and V4-Flash (284B, 13B), pretrained on 33T and 32T tokens, both natively support 1M context.

Efficiency at 1M tokens versus V3.2: V4-Pro needs 27% of single-token FLOPs and 10% of KV cache; V4-Flash needs 10% and 7%. The base V4-Flash (13B activated) beats V3.2-Base (37B activated) on most benchmarks with fewer parameters, especially world knowledge and long context.

The post-trained V4-Pro-Max sets a new open-source mark: SimpleQA 57.9 (about 20 points ahead of open models, trailing Gemini-3.1-Pro's 75.6), Chinese-SimpleQA 84.4, LiveCodeBench 93.5, Codeforces rating 3206 (above Gemini's 3052 and GPT-5.4's 3168, ranking 23rd among humans), which the authors call the first open model to match closed models in coding competitions. On MRCR at 1M tokens it beats Gemini-3.1-Pro and trails only Opus 4.6. Putnam-2025 reaches a proof-perfect 120/120.

Why it matters

DeepSeek is among the few in open source to actually make million-token context routine, and at lower compute and memory. That is infrastructure-grade good news for long-horizon agents, cross-document analysis, and test-time scaling. V4-Pro-Max pushes the open-source frontier on knowledge and reasoning, and matching closed models on coding competitions is a concrete step forward.

Limitations

The authors concede the architecture is complex (many validated components were retained to limit risk, and simplification is future work) and that Anticipatory Routing and SwiGLU Clamping work without being well understood. They place the reasoning capability about 3 to 6 months behind the strongest closed models, and agent tasks still trail closed models. This is a preview; multimodal and long multi-round agents are future work. Several benchmarks are run in an internal framework whose cross-model strictness is hard to verify externally, and FP4 times FP8 peaks match FP8 times FP8 on current hardware, so the further speedup awaits future hardware.

Terms

Source

What people are saying

Related papers

All paper explainers