New model adds engram and simplified mHC; community digs into KV cache tradeoffs

stochasticchasm · x · 2026-09-11

A researcher notes a new model incorporates the engram mechanism plus an mHC simplification — both expected additions — with mega-mhc flagged as worth investigating. Discussion also covers KV cache tradeoffs at 500K context: within a request the KV cache must live in HBM/RAM, but between requests, re-prefilling large chunks is negligible in FLOPS. Whether dspark was used in pre-training awaits the paper.

Related event: Inference-First Architecture Sparks Debate: FP4 KV Cache and Pure CSA2 Compression in Focus(11 posts)→

Original post →

More from Models

Models channel →