SemiAnalysis: Engram DRAM offloading delivers up to 50% better perf, upstreamed to vLLM
bookwormengr · x · 2026-09-23
SemiAnalysis published real-world numbers on DeepSeek's Engram design: offloading Engram to DRAM yields up to 50% better performance on H200, B200, B300 and even GB300 NVL72.
Key points:
- Engram is DeepSeek's approach to reducing HBM usage in favor of DRAM/LPDDR.
- SemiAnalysis worked with NVIDIA and AMD engineers to upstream ROCm Engram DRAM offloading support into vLLM (PR 57491), seeing up to 50% gains there too.
- Default recommended recipes are being updated to use Engram offloading via PRs 985, 1002, 1005 and 1006.
The poster quips that SemiAnalysis is god-tier at news reporting but slow as a forward-looking analyst, waking up to Engram late.
Related event: SemiAnalysis Tests Engram DRAM Offloading, Up to 50% Gains(2 posts)→
More from Infra
- Flash-dLLM accelerates diffusion LLMs up to 11x with IO-aware KV caching — MBZUAI · 2026-09-23
- vLLM v0.30.0 ships with 762 commits: watermarking, HiSparse, Model Runner V2 — vllm_project · 2026-09-23
- Sea first ASEAN company to adopt NVIDIA Vera Rubin as Nemotron spreads across Southeast Asia — NVIDIA Blog · 2026-09-23
- liuliu warns: claimed 4x-10x speedups over MLX or llama.cpp on Apple hardware are noise — teortaxesTex · 2026-09-23
- DeepSeek DSec cluster BoM estimated at ≤$20M; $1B could buy 50 clusters and 19M concurrent sandboxes — teortaxesTex · 2026-09-23
- B200 Rental Prices Hit All-Time High at $7.88 per GPU-Hour — sudoraohacker · 2026-09-23