GPUs idle waiting on memory while agent task lengths double every 4 months
alex_verem · x · 2026-10-07
The author argues the AI industry is built on the wrong infrastructure, citing Ben Horowitz and industry data:
- In chatbot workloads, a GPU spends only 5-15% of its time computing, with the rest waiting on memory — inference is bottlenecked by memory bandwidth, not compute.
- Nvidia's $20 billion deal for Groq's chips reflects this shift.
- Per Epoch AI, the length of tasks AI agents can perform doubles every 4 months, yet current systems were designed for millisecond-scale requests.
Core claim: infrastructure built for millisecond requests can't support long-running agent workloads, and the inference stack needs rebuilding for the agent era.
More from Infra
- 21M model + 6.4B SSD-resident lookup table matches a 114M dense model — fechyyy · 2026-10-07
- Ollama 0.35 adds local decision models from Cloudflare, Together and Bespoke — Technovangelist · 2026-10-07
- CostGraph becomes a drop-in InfraCost replacement for comparing GPU prices from L40 to A100 — saheedniyi_02 · 2026-10-07
- Oki Home launches a $1,799 Memory Computer running a 27B model locally with up to 16TB Memchip storage — Scobleizer · 2026-10-07
- AveniatsHub runs Qwen3 and Gemma3 fully offline on Android — MiraMooreIloveLLM · 2026-10-07
- Phonon-2 hits 606x real-time on a MacBook Air: one hour of speech in 6 seconds — julianweisser · 2026-10-07