Stanford and Together AI paper: hybrid local-cloud routing cuts AI cost and energy by 60-80%
rohanpaul_ai · x · 2026-09-11
A new Stanford/Together AI paper, Intelligence per Watt: Measuring Intelligence Efficiency of Local AI, benchmarks local AI efficiency:
- Hybrid local-cloud routing cut energy, compute and cost by 60-80% vs. a batched-cloud baseline.
- From 2023 to 2025, local AI's intelligence-per-watt improved 5.3x, with locally serviceable query coverage rising from 23.2% to 71.3%.
- An iPhone 16 Pro delivered 7x higher intelligence-per-watt than workstation GPUs on the same model and precision.
- A pool of 20+ local models, routed per query, beat 3 frontier cloud models on 3 of 4 benchmarks.
- Dropping FP16 to FP4 cut inference energy 3-3.5x at 2.5 accuracy points per precision step.
- The hard end remains weak: 95% of the hardest reasoning problems were still unsolved locally.
Related event: Stanford Paper: Hybrid Local-Cloud Routing Cuts 60-80% of AI Costs(2 posts)→
More from Infra
- Spain's hourly 80% renewable matching rules clash as France fast-tracks 700MW sites, UK cuts grid queues — eherrerosj · 2026-09-11
- AI could add 0.3-0.4 points to Europe's productivity growth, but the EU holds under 5% of global compute — rohanpaul_ai · 2026-09-11
- Qualcomm's next-gen Hexagon NPU runs 30B MoE models with 32K context on-device — lee_stott · 2026-09-11
- Rented GPU Bills: Host CPU and Script Defaults Made Costs 31x Higher — Worldly_North_7213 · 2026-09-11
- Routing NVIDIA PAIR to llama.cpp on an AMD ROCm node (2×R9700): full notes — Don_Reuter · 2026-09-11
- Running Qwen3.8 locally on a 128GB laptop for agentic coding: thinking tokens, not tok/s, set the wall clock — deepu105 · 2026-09-11