Local AI models improve 18x in accuracy per joule over 16 months, Stanford team says

Azaliamirh · x · 2026-09-12

Inference is going increasingly hybrid: a Stanford-affiliated team reports that accuracy per joule of local models improved 18x in just 16 months — 5.9x from hardware gains and 3.0x from model improvements. The work, by Azaliamirh, Avanika, Jon Saad-Falcon, Hazy Research and John Hennessy, was featured in a Financial Times deep-dive on the world's decreasing dependence on centralized cloud AI.

Original post →

More from Infra

Infra channel →