DeepSeek V4.1 Flash's Engram memory layer trades FFN compute for lookup tables, SemiAnalysis data suggests
teortaxesTex · x · 2026-09-26
A widely shared breakdown (via SemiAnalysis data, unverified by DeepSeek) claims V4.1 Flash's Engram memory layer converts common phrases into lookup tables: only 0.8B/1.6B params active per layer, with 196B as lookup tables. Moving tables to host memory on B300 cuts TP4 to TP2 (1.6x cost-efficiency), boosts KV capacity 36% on 4x GB300, and beats SSD offloading on B200 (121M vs 52M tokens per dollar). Dual GB300 nodes hit 56K tok/s prefill and 253 tok/s single-user decode. The takeaway: cheapness comes from less compute and less HBM pressure — memory bandwidth matters more than capacity. Critics note cheapness doesn't prove modeling gains.
More from Infra
- Anthropic and OpenAI list identical headline prices, but cache read differs 4x — julsimon · 2026-09-26
- ASML filings show EMEA sales collapsing to zero as Europe builds 2015-era chips — julsimon · 2026-09-26
- Glamsterdam will replace sync healing with state diffs, further speeding up Ethereum nodes — banteg · 2026-09-26
- Unverified claim: xAI's 200k GB300 cluster at 10% MFU sparks community pushback — teortaxesTex · 2026-09-26
- NVIDIA at $5.4 trillion is now worth more than the entire UK or French stock market — iamfakhrealam · 2026-09-26
- A Curated Paper List for Learning Distributed LLM Training and Inference — East-Muffin-6472 · 2026-09-26