DeepSeek slashes KV cache to 890 bytes/token, hinting at million-agent swarms

teortaxesTex · x · 2026-09-10

Analysis of DeepSeek V4.1 Flash's engineering: KV cache compressed to 890 bytes/token (<1GB per 1M tokens), cutting HBM requirements 3.9x and SSD needs 8x versus V4-Flash — half a year on, no other model approaches V4 arch's KV efficiency. On a 96GB/950DT minimal setup, the author estimates 1.3 million max-budget agents could run in parallel on the Huawei cluster acquired in May. Margins reportedly exceed 90%; 'if this were an American lab, we'd be talking about RSI.'

Related event: Leaked DeepSeek V4.1 Benchmarks Point to New 552B Architecture(15 posts)→

Original post →

More from Infra

Infra channel →