DeepSeek V4.1 Flash's Engram memory layer trades FFN compute for lookup tables, SemiAnalysis data suggests

teortaxesTex · x · 2026-09-26

A widely shared breakdown (via SemiAnalysis data, unverified by DeepSeek) claims V4.1 Flash's Engram memory layer converts common phrases into lookup tables: only 0.8B/1.6B params active per layer, with 196B as lookup tables. Moving tables to host memory on B300 cuts TP4 to TP2 (1.6x cost-efficiency), boosts KV capacity 36% on 4x GB300, and beats SSD offloading on B200 (121M vs 52M tokens per dollar). Dual GB300 nodes hit 56K tok/s prefill and 253 tok/s single-user decode. The takeaway: cheapness comes from less compute and less HBM pressure — memory bandwidth matters more than capacity. Critics note cheapness doesn't prove modeling gains.

Original post →

More from Infra

Infra channel →