Entrepreneur calls current AI inference 'idiotically' inefficient; in-place weight architectures 1000x better

whurley · x · 2026-09-20

A retweeted take from Dave Blundin argues current AI inference is deeply wasteful: weights are pulled from HBM into a GPU, used for a picosecond, and dumped, over and over. We run AI on chips designed for graphics, and the market hasn't priced that in. Architectures that keep weights in place could be 1,000–1,000,000x more efficient, with the only blocker being a supply chain that can't keep up with fast-evolving model architectures. He discussed this on stage with Andrew Feldman, Atiq Raza and John Werner.

Original post →

More from Infra

Infra channel →