Running a 177B MoE Model on Dual 3090s and an Old EPYC: 38 tok/s for $800

ludos1978 · reddit · 2026-09-13

The OP shares their hardware setup and benchmarks for running Qwen3-Flash-Next locally (177B total params / 6B active MoE, IQ4XS quantization):

With $800 left in the budget, the OP is torn between adding a third 3090 or upgrading to a Zen2 EPYC (drop-in on the same socket), and asks whether more/faster RAM would actually help—the goal is keeping speeds up during parallel multi-agent workloads.

Original post →

More from Infra

Infra channel →