DeepSeek's Engram and Linear Attention: Trading FLOPs for HBM in Inference
bookwormengr · x · 2026-09-02
A discussion highlights the emerging wave of algorithmic and architectural optimizations that trade more FLOPs for less memory in AI inference. While recomputing activations during the backward pass is known in training, inference is seeing innovations like DeepSeek's Engram, which hosts large embedding tables in LPDDR to save HBM and FLOPs—adopted by Longcat 2 and Qwen-3-Next. Additionally, linear attention mechanisms (as in Kimi K3, GLM-5.3-Flash) that leverage available SRAM are expected to work wonders for inference performance.
Related event: Trading FLOPs for HBM memory emerges as inference optimization trend(3 posts)→
More from Infra
- Not Diamond releases model routing method, cuts costs 20-80% — rohanpaul_ai · 2026-09-02
- Trading VRAM for FLOPs: Local Model Quantization Debate — max_paperclips · 2026-09-02
- DGX Spark Cluster vs AMD Epyc Server: A Cost-Benefit Comparison — LeftHandHaku · 2026-09-02
- User complains about dev tool hogging 98% RAM, requests local GPU support — MickeySteamboat · 2026-09-02
- Nvidia Backstop Enables $947M in Bank Loans for GPU Purchases — rohanpaul_ai · 2026-09-02
- Dell Predicts 87x Surge in AI Inference Demand by 2030, Enterprise Agents to Dominate Workloads — toptickcrypto · 2026-09-02