Leak: DeepSeek V4.1 reportedly drops V4's HCA, generalizes CSA, removes sparse-attention warmup

zephyr_z9 · x · 2026-09-10

KOL teortaxesTex claims DeepSeek V4.1 strips out V4's workarounds at every level: HCA is ditched, CSA is generalized, no sparse-attention warmup, a simpler single-pass mHC, plus Engram, calling it a 'new Transformer'. He earlier predicted DeepSeek would compress V4's architecture into a simpler set of primitives, arguing DSA isn't the limit and only real architectural exploration crosses deep valleys. Unverified rumor.

Related event: Leaked DeepSeek V4.1 Benchmarks Point to New 552B Architecture(15 posts)→

Original post →

More from Models

Models channel →