DeepSeek's Engram architecture adopted by Qwen and Longcut to save HBM
bookwormengr · x · 2026-09-02
A discussion on inference optimization highlights the adoption of DeepSeek's Engram architecture by models like Longcat 2 (1.6T) and Qwen-3-Next. This technique allows hosting larger embedding tables in LPDDR, reducing HBM storage requirements and model parameters, thus lowering FLOPs. Coupled with innovations like Linear Attention that leverage available SRAM, these architectural changes are expected to significantly boost inference efficiency in the coming days.
Related event: Trading FLOPs for HBM memory emerges as inference optimization trend(3 posts)→
More from Infra
- AI Kernel Gen is Easy; End-to-End Stack Enablement is the Real Metric — clattner_llvm · 2026-09-02
- CMP 170HX Failures: 2 Dead in 2 Weeks, Defective Cores — cantgetthistowork · 2026-09-02
- Is it worth switching to AMD for AI Video Generation? ROCm Inquiry — iridescentblob · 2026-09-02
- HybridInfer: Router auto-falls back to cloud when local model wedges — simrankoulsm · 2026-09-02
- Looped transformer rumors may be true, bearish for HBM demand — zephyr_z9 · 2026-09-02
- Buying GPUs to run local models is pointless, argue devs: cloud APIs win on cost — brandon_galang · 2026-09-02