SemiAnalysis: Why GLM-5.3 Sparse Attention Doesn't Cut HBM Memory Capacity Needs
burny_tech · x · 2026-09-29
SemiAnalysis published a deep dive into how GLM-5.3's sparse attention affects DRAM/HBM memory economics.
- Sparse attention selects only top-k relevant tokens for SDPA, cutting KV cache bandwidth needs — but top-k selection typically requires the full context in HBM, so it doesn't remove the memory capacity bottleneck, and throughput remains capacity-limited (per HiSparse).
- The SGLang team built HiSparse, a hierarchical memory system that proactively offloads KV cache entries from HBM to host DRAM with LRU-style eviction; to hide miss latency, it overlaps layer-N KV loading with layer N-1 execution (building on HiCache).
- Result: major throughput gains at high concurrency and long contexts, at the cost of top-k cache-miss I/O overhead.
- The piece also covers DeepSeek Sparse Attention, IndexShare, single-rollout async optimization, AgentX TileRT, and implications for HBM/NAND memory TAM (paid).
More from Infra
- Only defense contractors have wafer-scale GaN HEMT + InP HBT on CMOS, DARPA-funded — jwt0625 · 2026-09-29
- AMD reportedly already selling 2028 CPU volume; Venice yield issues less severe than rumored — Sethwinterroth · 2026-09-29
- Brookings paper projects $10.3T in AI infrastructure investment by 2032, flags hidden financing risks — VraserX · 2026-09-29
- Training-free WaveFront Decoding speeds up looped LMs by up to 4.81x — SNU-VLSI · 2026-09-29
- Exploit Summit Montreal recap: Gamma tokens, iota SDK, $12M run rate for Targon — markjeffrey · 2026-09-29
- Bain says AI must earn $6T a year by 2031 — matching all global IT spending today — sanjaykalra · 2026-09-29