Stanford's Intelligence per Watt: hybrid local-cloud routing cuts cost 60-80%
rohanpaul_ai · x · 2026-09-11
A new Stanford/Together AI paper (authors incl. Christopher Ré, John Hennessy) proposes Intelligence per Watt (IPW) — task accuracy per unit of power — as a unified metric for local LLM inference.
- Evaluated 20+ local SOTA models (≤20B active params), 8 accelerators, and 1M real-world chat/reasoning queries.
- Hybrid local-cloud routing cut energy, compute, and cost by 60–80% vs a batched-cloud baseline.
- From 2023 to 2025, local AI's intelligence-per-watt improved 5.3×; locally serviceable query coverage jumped from 23.2% to 71%+.
The takeaway: small local models plus laptop-class accelerators (e.g. Apple M4 Max) can meaningfully redistribute demand away from centralized cloud infrastructure.
Related event: Stanford Paper: Hybrid Local-Cloud Routing Cuts 60-80% of AI Costs(2 posts)→
More from Infra
- DeepSeek launches V4.1-Flash with 1M-token context and 4x smaller KV-cache — matlabulous · 2026-09-11
- What Can You Still Run on 8GB VRAM? User Asks for Small Models With Tool Use — riceinmybelly · 2026-09-11
- Spain's hourly 80% renewable matching rules clash as France fast-tracks 700MW sites, UK cuts grid queues — eherrerosj · 2026-09-11
- AI could add 0.3-0.4 points to Europe's productivity growth, but the EU holds under 5% of global compute — rohanpaul_ai · 2026-09-11
- Qualcomm's next-gen Hexagon NPU runs 30B MoE models with 32K context on-device — lee_stott · 2026-09-11
- Stanford and Together AI paper: hybrid local-cloud routing cuts AI cost and energy by 60-80% — rohanpaul_ai · 2026-09-11