What to Run at 128GB VRAM? A Builder Weighs Quantized Qwen, GLM, and DeepSeek
Madigan37 · reddit · 2026-09-10
A hobbyist upgrading to 128GB VRAM asks what to run locally. He was leaning toward Qwen3.8 Flash-Next at Q4 (preferring nothing below Q4), but the lukewarm reception to Flash-Next has him eyeing GLM 5.3 at Q2 or DeepSeek 4 Flash at Q2/Q3 instead. He plans to try all three and wants to hear what people in the same boat are running.
More from Infra
- Autonomous adds Omarchy OS to its $26,100 dual-RTX-5090 AI workstation — dee_hw · 2026-09-10
- 100M output tokens for $60: DeepSeek off-peak pricing undercuts Opus 5 by 40x — airesearch12 · 2026-09-10
- Dev take: token demand will grow far faster than demand for top-line intelligence — willcb · 2026-09-10
- Arm lands Lenovo and ByteDance's Volcengine as first China customers for its AI server chips — pstAsiatech · 2026-09-10
- d-Matrix adopts NVIDIA NVLink Fusion to bring Raptor XPUs to rack-scale deployment — nordicinst · 2026-09-10
- Why DeepSeek might profit despite open weights: it's the only one happy to optimize for its own architecture — yacineMTB · 2026-09-10