DeepSeek's Engram architecture adopted by Qwen and Longcut to save HBM

bookwormengr · x · 2026-09-02

A discussion on inference optimization highlights the adoption of DeepSeek's Engram architecture by models like Longcat 2 (1.6T) and Qwen-3-Next. This technique allows hosting larger embedding tables in LPDDR, reducing HBM storage requirements and model parameters, thus lowering FLOPs. Coupled with innovations like Linear Attention that leverage available SRAM, these architectural changes are expected to significantly boost inference efficiency in the coming days.

Related event: Trading FLOPs for HBM memory emerges as inference optimization trend(3 posts)→

Original post →

More from Infra

Infra channel →