GLM 5.3 and Qwen 3.8 now run really well locally on single desktops
jasonkneen · x · 2026-09-07
Amid the latest AGI debates, the author points to something more concrete: local models are now genuinely good. GLM 5.3 and Qwen 3.8 can run on a single box or dual NVIDIA Sparks, meaning capable open-weight models are now a practical desktop option for everyday developers — arguably the most important story being overlooked.
More from Infra
- Tracking one tennis ball with GPT-6 burns 7.87M tokens — Scobleizer · 2026-09-07
- Huawei's Kirin 9050Pro uses logic folding to cut NPU power 66%, run 30B MoE on-device — APPSO · 2026-09-07
- Energy and the power grid, not chips, are the real bottleneck for AI at 400k-GPU scale — kimmonismus · 2026-09-07
- Signing TLS handshakes inside a TPM to protect machine identity — jedisct1 · 2026-09-07
- Same RTX 5080, 70% slower: renting GPUs across regions can silently cost you speed — Grouchy_Television75 · 2026-09-07
- Adding an RTX 2000 Ada 16GB (75W) to a gaming PC for local LLM inference under a 700W PSU — tableball35 · 2026-09-07