GPUs idle 85-95% waiting on memory: why the whole AI stack is being rebuilt
alex_verem · x · 2026-10-07
- Citing Ben Horowitz and several data points, the author argues today's AI infrastructure is mismatched: a GPU serving a chatbot spends only 5–15% of its time computing, the rest waiting on memory—why Nvidia reportedly paid $20B for Groq's chips (unverified).
- Key data: the task length AI agents can handle doubles every four months, yet systems were built for millisecond requests; Epoch AI projects public internet training data runs out around 2028.
- The rebuild: AI-only chips, datacenters for day-long jobs, models learning from the real world. a16z just raised $2.8B to bet on exactly that.
More from Infra
- Mistral CEO: Large 4 trained on our own compute, 'RL shows no sign of saturation' — sivareddyg · 2026-10-07
- HF engineer releases open slide deck on local AI: quantization to speculative decoding — mervenoyann · 2026-10-07
- 21M model + 6.4B SSD-resident lookup table matches a 114M dense model — fechyyy · 2026-10-07
- Ollama 0.35 adds local decision models from Cloudflare, Together and Bespoke — Technovangelist · 2026-10-07
- CostGraph becomes a drop-in InfraCost replacement for comparing GPU prices from L40 to A100 — saheedniyi_02 · 2026-10-07
- Oki Home launches a $1,799 Memory Computer running a 27B model locally with up to 16TB Memchip storage — Scobleizer · 2026-10-07