Entrepreneur calls current AI inference 'idiotically' inefficient; in-place weight architectures 1000x better
whurley · x · 2026-09-20
A retweeted take from Dave Blundin argues current AI inference is deeply wasteful: weights are pulled from HBM into a GPU, used for a picosecond, and dumped, over and over. We run AI on chips designed for graphics, and the market hasn't priced that in. Architectures that keep weights in place could be 1,000–1,000,000x more efficient, with the only blocker being a supply chain that can't keep up with fast-evolving model architectures. He discussed this on stage with Andrew Feldman, Atiq Raza and John Werner.
More from Infra
- Jevons paradox is classic low-end disruption you won't spot from a GPU-rich hyperlab — cramforce · 2026-09-20
- SpaceX's orbital AI data centers weigh up to 4,000 kg each, filing seeks 1M satellites — XFreeze · 2026-09-20
- Local Models for Personal Agents: GPT Luna Surprises a Coding-Agent Veteran — gized00 · 2026-09-20
- OpenAI hardware VP details first custom chip Jalapeño and its nine-month tape-out — bigdata · 2026-09-20
- focus-llama: a llama.cpp fork implementing Declarative Attention for up to 0.71x decode time — Ok-Shower7286 · 2026-09-20
- VTrain, a Vulkan-based resident trainer, fixes memory leak and offloads more work to GPU — Savantskie1 · 2026-09-20