Six efficiency breakthroughs labs didn't see coming upend semiconductor demand assumptions
bookwormengr · x · 2026-09-16
An engineer lists six trends most AI lab insiders failed to foresee: 4-hi HBM (equivalent HBM reduction via WideEP and smaller batches); LPDDR-based approaches cutting parameters in HBM and FLOPs by 50%; hybrid linear and sparse attention proving highly effective; Flash-style models excelling at economically valuable tasks beyond coding; techniques like Attention Residue, mHC and stable MoE packing denser intelligence per parameter; and 400X KV cache compression via encoder-decoder with little quality loss. Implication: most labs assumed near-linearly growing model sizes (1-3-5-10-30-50-100T), which isn't happening — yet most long-term agreements were signed on that old assumption.
More from Infra
- Running Qwen 27B and DeepSeek v4 Flash together on one heterogeneous machine — samsja19 · 2026-09-16
- JPMorgan sees 25M+ GPU/ASIC shipments by 2028, ASICs dominate — a 'narrative violation' — bookwormengr · 2026-09-16
- iamtrask: The Endgame Is a Trust Web of Personal LLM Servers, Not One AGI — iamtrask · 2026-09-16
- Program-as-Weights: 0.6B interpreter matches Qwen3-32B prompting with 1/50 the memory — yuntiandeng · 2026-09-16
- Intel bets on CPU-side KV Cache offloading and QAT hardware compression for agent inference — 量子位 · 2026-09-16
- Trillion-parameter model trained on own physics lab data hits memory-efficiency SOTA — vwxyzjn · 2026-09-16