Benchmarking 2.4T Qwen Model on Local GPUs
Developers tested the 2.4T Qwen3.8-A95B model on local consumer GPUs, utilizing five graphics cards including RTX 5090s. The benchmark revealed a generation speed of about 0.8 tokens per second, highlighting the extreme hardware demands of massive local models.
2026-08-14 ~ 2026-08-14 · 2 related posts
- Running 2.4T Qwen3.8 Model on RTX 5090 + 5060 Ti: 0.8 tok/s Tested — mossy_troll_84 · 2026-08-14
- Running Qwen 2.4T Locally: 5 GPUs Still Can't Make It Viable — klicker0 · 2026-08-14