DeepSeek's Engram and Linear Attention: Trading FLOPs for HBM in Inference

bookwormengr · x · 2026-09-02

A discussion highlights the emerging wave of algorithmic and architectural optimizations that trade more FLOPs for less memory in AI inference. While recomputing activations during the backward pass is known in training, inference is seeing innovations like DeepSeek's Engram, which hosts large embedding tables in LPDDR to save HBM and FLOPs—adopted by Longcat 2 and Qwen-3-Next. Additionally, linear attention mechanisms (as in Kimi K3, GLM-5.3-Flash) that leverage available SRAM are expected to work wonders for inference performance.

Related event: Trading FLOPs for HBM memory emerges as inference optimization trend(3 posts)→

Original post →

More from Infra

Infra channel →