Running a 180B MoE (5B active) on one RTX 3090 + 128GB RAM: full config and 15.5 tok/s benchmarks

cezarducatti · reddit · 2026-09-05

A developer shares a complete llama.cpp setup and real benchmarks for running Qwen3.8-Flash-Next (UD-Q4KXL, 180B total / 5B active MoE) on a single RTX 3090 24GB plus 128GB DDR4, and asks for tuning advice:

Original post →

More from Infra

Infra channel →