Running Qwen 2.4T Locally: 5 GPUs Still Can't Make It Viable
klicker0 · reddit · 2026-08-14
A hardcore user tested running the massive 2.4T parameter Qwen3.8 model locally using an extreme UD-Q10 quantized version (397GB).
- Hardware: Dual RTX 5090s + three RTX 3090s + 96GB system RAM. Even with this setup, the model's requirements exceeded the total available VRAM and RAM by double.
- Performance: The generation speed was a mere 0.25 tokens/sec (taking 11 minutes and 38 seconds for 178 tokens).
- Conclusion: This speed is practically unusable, even for overnight tasks. While having open weights for massive models is great, consumer hardware still falls far short of running a 2.4T model locally.
Related event: Benchmarking 2.4T Qwen Model on Local GPUs(2 posts)→
More from Infra
- Inference Engineering for DeepSeek V4 Pro 0813: A 1.7T Open Model — philipkiely · 2026-08-14
- NVIDIA Driver Update Silently Enables ECC, Costing Consumer GPUs 1.5GB VRAM — MastMaithun · 2026-08-14
- High Bandwidth Flash Could Break the VRAM Bottleneck: An AGI for $40k? — Zombiecidialfreak · 2026-08-14
- Alibaba Open-Sources MNN: A Lightweight Inference Engine for On-Device LLMs — tom_doerr · 2026-08-14
- AMD Gifting Ryzen AI Halo Box Optimized for LLM Workloads — JosephJacks_ · 2026-08-14
- How a GPU Actually Works: The Intuition LLM Engineers Need — burny_tech · 2026-08-14