Running 2.4T Qwen3.8 Model on RTX 5090 + 5060 Ti: 0.8 tok/s Tested
mossy_troll_84 · reddit · 2026-08-14
A developer successfully ran the massive 2.4T parameter Qwen3.8-2.4T-A95B model on consumer hardware, sharing detailed benchmarks and configuration insights.
Hardware & Environment
- GPUs: RTX 5090 (32GB) + RTX 5060 Ti (16GB)
- Other: AMD Ryzen 9 9950X3D, 128GB DDR5 RAM
- Model: Unsloth Q10 GGUF quantization (397 GiB)
Performance
In a controlled 32-token generation test, the setup achieved a generation speed of 0.80 tok/s. Enabling native MTP (speculative decoding) resulted in a 3.64% throughput increase and a 2.74% reduction in total wall time, with an MTP acceptance rate of 90.48%.
Configuration
The author shared optimal llama.cpp parameters, including a tensor split of 4,1 and offloading 91 MOE experts to the CPU. To prevent MTP from running out of VRAM, block 92 expert tensors were forced to remain on the CPU.
More from Infra
- Databricks Introduces Smart Routing in Unity AI Gateway, Claims 30%+ Cost Reduction — matei_zaharia · 2026-08-14
- OpenAI Acquired 4.2% Stake in Cerebras Before Ultrafast Launch — ryanmerket · 2026-08-14
- Jensen Huang Marks DGX 10th Anniversary, Unveils DGX Spark at 5x Original Power — nvidia · 2026-08-14
- NVIDIA x Runway: Gen-4.5 Integrated into Vera Rubin Platform in One Day — nvidia · 2026-08-14
- RTX 5090 Test: SageAttention Nearly Doubles Video Generation Speed — gabxav · 2026-08-14
- CoreWeave Sandbox Lets You Drive Claude Agents From Your Phone — wandb · 2026-08-14