With 8-10T parameter models coming next year, who can still run frontier LLMs locally?

power97992 · reddit · 2026-10-01

A Reddit discussion on local runnability of next-gen frontier models: DeepSeek has said it will release an 8T model, Qwen plans a 10T one, and Kimi will likely follow, with flash models around 1-2T parameters.

The author estimates running a 4.4-bit 8T model with full context would take roughly nine 512GB M5 Ultras or 48 RTX 6000 Pros. Conclusion: only companies, cloud providers and the wealthy will run pro models locally; most people will struggle even with a 1T flash model and will likely settle for smaller models like Qwen 5 27B or the cloud.

Original post →

More from Infra

Infra channel →