ENGRAM: 50% Higher Revenue per GW, But Is It Really a Free Lunch for AI Models?
bookwormengr · x · 2026-09-27
A deep dive into ENGRAM, the technique that caught the industry's attention after SemiAnalysis claimed it could deliver 50% higher revenue per GW — worth many billions.
- How it works: stores sequence embeddings (meanings of token sequences) in embedding tables that can be offloaded to DRAM, saving costly HBM
- Benefits: frees space for KV Cache and reduces FLOPs per token, cutting memory and accelerator requirements
- Results: dramatic performance gains per SemiAnalysis, which even published a Dylan Patel roast acknowledging earlier analysts were wrong
- The author questions whether it's truly a free lunch and whether it impacts model quality, referencing their own coverage since Feb 2026
More from Infra
- DRAM output ~50EB vs NAND ~1100EB this year; CXL pooling is about utilization, not replacing eSSD — zephyr_z9 · 2026-09-27
- Memory Attention: Paper Replaces Value Projection With Lookup, Cuts GPU Storage 7.38% — serrjoa · 2026-09-27
- Bought a 64GB DDR5 kit 3 years ago for 17k rupees — now it's worth 90k as RAM prices soar — ojasvi_yadav · 2026-09-27
- Thermocompression bonding a 200mm wafer: 70 kN at 400-450°C for 20-45 min — jwt0625 · 2026-09-27
- OpenAI reportedly facing ugly compute shortage as Pro quotas become the new Plus — ns123abc · 2026-09-27
- UK village revolts against 1.5GW AI datacentre planned in UNESCO biosphere reserve — nordicinst · 2026-09-27