Local AI models improve 18x in accuracy per joule over 16 months, Stanford team says
Azaliamirh · x · 2026-09-12
Inference is going increasingly hybrid: a Stanford-affiliated team reports that accuracy per joule of local models improved 18x in just 16 months — 5.9x from hardware gains and 3.0x from model improvements. The work, by Azaliamirh, Avanika, Jon Saad-Falcon, Hazy Research and John Hennessy, was featured in a Financial Times deep-dive on the world's decreasing dependence on centralized cloud AI.
More from Infra
- DeepSeek v4.1 Flash runs out of the box on six NVIDIA GPUs via vLLM on day 0, AMD lags — woosuk_k · 2026-09-12
- Open-source Plano routes LLM calls by prompt intent, no agent code changes, cutting bills 2x — Roger_M_Taylor · 2026-09-12
- Netflix Engineer Open-Sources Headroom, Cuts Agent Token Use by Up to 95% — Roger_M_Taylor · 2026-09-12
- DeepSeek V4.1 Flash cuts global KV cache to 890 bytes/token, but HBM demand may rise with agent swarms — teortaxesTex · 2026-09-12
- Polymarket Prices AI Bubble Burst at 12% Odds Through End of 2026 — Polymarket · 2026-09-12
- SGLang hits 873 tok/s on DeepSeek V4.1 Flash within 24 hours of launch — BanghuaZ · 2026-09-12