Running a 177B MoE Model on Dual 3090s and an Old EPYC: 38 tok/s for $800
ludos1978 · reddit · 2026-09-13
The OP shares their hardware setup and benchmarks for running Qwen3-Flash-Next locally (177B total params / 6B active MoE, IQ4XS quantization):
- Setup: EPYC 7551 (Zen1, 32c/64T) + Supermicro H11SSL-i + 128GB DDR4-2133 in an eight-channel config + 2×RTX 3090 (48GB) + a 1500W PSU
- Approach: llama.cpp, with expert weights kept in system RAM and hot experts cached in VRAM
- Performance: 38 tok/s on a single request; two parallel requests each drop to 4 tok/s
With $800 left in the budget, the OP is torn between adding a third 3090 or upgrading to a Zen2 EPYC (drop-in on the same socket), and asks whether more/faster RAM would actually help—the goal is keeping speeds up during parallel multi-agent workloads.
More from Infra
- Team With 2x A100 Grant Open-Sources Roadmap for Distilling Offline Edge Models — Lerok-Persea · 2026-09-13
- Prediction: hyperscalers will tighten AI capex in 2027 — pdamodaran · 2026-09-13
- llama.cpp launches llama.app: frontier AI fully local with one command — ngxson · 2026-09-13
- Qwen on M2 Ultra: latest oMLX update brings substantial local inference speedup — Thrumpwart · 2026-09-13
- Google rewrites SERP links to opaque google.com/goto redirects, raising the cost of scraping results at scale — jedisct1 · 2026-09-13
- Epoch: Huawei can't catch Nvidia this decade, GPT-6 Astra sets records, data center power doubles every 10 months — Epoch AI · 2026-09-13