Local models will always cost more — they win only when models get cheap
kuza55 · x · 2026-09-03
In a debate with GergelyOrosz, kuza55 argues local models (already real on phones) will always be more expensive than cloud because servers get far better utilization — so local wins only when the models you want are cheap, not when they're expensive. Analogy: gaming needs local latency for rendering, yet every multiplayer game still runs on rented cloud servers, since nearly everything has moved to SaaS.
Related event: Debate: Local Compute vs. Cloud — Local Models Are Always Pricier(2 posts)→
More from Infra
- Reddit: Picking the Best Chat Model for a 3090 Ti Local AI Butler — MarcusAurelius68 · 2026-09-03
- MTP vs MTP+Ngram on Qwen3.8 Flash: 10% Speed but 3x Token Usage — esw123 · 2026-09-03
- Perplexity Computer demos fully local operation on NVIDIA DGX Spark — chrmanning · 2026-09-03
- llama.cpp deprecates --chat-template-kwargs, reasoning-preserve now on by default — Bulky-Priority6824 · 2026-09-03
- Agentic API adds a stateful layer in front of vLLM for open-model agent runtimes — techNmak · 2026-09-03
- Google's Gemini 3.8 Flash 'works harder' but may burn more tokens at same pricing — The Verge AI · 2026-09-03