With 8-10T parameter models coming next year, who can still run frontier LLMs locally?
power97992 · reddit · 2026-10-01
A Reddit discussion on local runnability of next-gen frontier models: DeepSeek has said it will release an 8T model, Qwen plans a 10T one, and Kimi will likely follow, with flash models around 1-2T parameters.
The author estimates running a 4.4-bit 8T model with full context would take roughly nine 512GB M5 Ultras or 48 RTX 6000 Pros. Conclusion: only companies, cloud providers and the wealthy will run pro models locally; most people will struggle even with a 1T flash model and will likely settle for smaller models like Qwen 5 27B or the cloud.
More from Infra
- Micron FQ4 Revenue Hits $54.2B, Up 379% YoY; Guides Record $61.5B Next Quarter — Beth_Kindig · 2026-10-01
- Azure AI kicks off series: why content extraction matters more as GenAI models get stronger — adnan_hashmi · 2026-10-01
- Micron Crushes Estimates as Quarterly Revenue Nearly Quadruples to $54.2B — Polymarket · 2026-10-01
- Google's compute hunger games: sells 1M TPUs to Anthropic while renting SpaceX Blackwell at 2x market — zacharynado · 2026-10-01
- Micron plans to return 100% of excess cash to shareholders from Dec 2026 — firstadopter · 2026-10-01
- Ex-OpenAI policy VP: compute is now split into ~3 tiers and the gap is widening — Miles_Brundage · 2026-10-01