New model adds engram and simplified mHC; community digs into KV cache tradeoffs
stochasticchasm · x · 2026-09-11
A researcher notes a new model incorporates the engram mechanism plus an mHC simplification — both expected additions — with mega-mhc flagged as worth investigating. Discussion also covers KV cache tradeoffs at 500K context: within a request the KV cache must live in HBM/RAM, but between requests, re-prefilling large chunks is negligible in FLOPS. Whether dspark was used in pre-training awaits the paper.
More from Models
- Microsoft Patches Record 974 Vulnerabilities, Mostly Found by AI — Distinct-Question-16 · 2026-09-11
- A Four-Step Verification Method to Catch AI That Fakes Reading Financial Reports — anthara_ai · 2026-09-11
- DeepSeek V4.1 Flash tops Vals open-weight index at $0.30 per test, with the smallest skills gap — teortaxesTex · 2026-09-11
- Do You Really Need Flagship Models? Dev Argues Medium Effort Covers 80% of Coding — iamaliveix · 2026-09-11
- OpenAI appears to be quietly rolling out managed Agents on its platform — testingcatalog · 2026-09-11
- 30B Open Model OpenResearcher Beats GPT-4.1 on BrowseComp-Plus — TheZachMueller · 2026-09-11