Offloading Engram to DRAM boosts H200/B200 performance by up to 50%
zephyr_z9 · x · 2026-09-21
SemiAnalysis reports that offloading Engram to DRAM yields up to 50% better performance on H200, B200, B300, and GB300 NVL72. Its team, working with engineers from Pi, also upstreamed ROCm Engram DRAM offloading support to vLLM (PR 57491) with similar gains. PRs 985, 1002, 1005, and 1006 are updating default recommended recipes from NVIDIA, AMD, and SemiAnalysis to enable Engram offloading.
More from Infra
- OpenAI removes Ultrafast tier from GPT-5.6 Sol in Codex, fueling GPT-6 Sol rumors — imjustnewatai · 2026-09-22
- Is upgrading from 2x to 4x RTX 3090 worth it for local LLM work? — fgoricha · 2026-09-22
- Qwen-Image local on a 24GB MacBook Pro takes 5-6 minutes per image — vista8 · 2026-09-22
- Cerebras CEO on Jensen Huang: a decade trading as 'nobody' before Nvidia made it — rohanpaul_ai · 2026-09-22
- How Tencent Hunyuan packed a 770B model into 214 GiB with 5-bit-per-4-weights quantization — TencentHunyuan · 2026-09-22
- 456GB DeepSeek v4.1 runs locally at 40 tok/s with Threadripper + dual RTX 6000 hybrid setup — HankYeomans · 2026-09-22