DeepSeek V4.1 Flash tech report: KV cache compression lets 552B beat 1.6T

机器之心 · wechat · 2026-09-12

Machine Heart's deep dive into the DeepSeek-V4.1-Flash tech report: a 552B-backbone + 196B Engram multimodal MoE that beats the 1.6T-parameter V4 Pro (scoring 40 on ArtificialAnalysis) by compressing KV cache to the extreme — global KV down to 890 bytes/token, 437x less than V1.

Key techniques:

Other highlights: Single-Pass mHC halves activation memory traffic; the Engram N-gram-hash memory module ships in a real model for the first time, offloading static knowledge to host memory; the DSpark speculative decoding module is trained post-pretraining and accelerates both serving and RL rollouts. Post-training openly has "no algorithmic novelty," but introduces a 1–100 reasoning effort scalar: raising effort from 25 to 100 lifts average Pass@1 across eight reasoning benchmarks from 67.1% to 76.3% at 2.5x output tokens; API tiers max/high/low map to b=100/75/50. Context scaling from 4K to 1M adds only 1/4 decode FLOPs per token; despite verbosity (89k tokens/task) it costs just $0.27 per task.

Original post →

More from Infra

Infra channel →