Running Qwen Flash Next on a $1k RAM-rich GPU-poor box: 12tps but far better output

Positive-Stock6444 · reddit · 2026-08-30

A Reddit user shares benchmarks from running Qwen models locally on a 2018 ThinkStation P520 (Xeon W-2145, 256GB quad-channel DDR4, 12GB 3060, €1,000 total).

Qwen3.6 35B A3B Q4KM (former daily driver): 20GB RAM resident, 400tps prefill, 30-50tps generation, 128k context, and tolerant of other load on the box.

Qwen3.8 Flash Next UD-Q4KXL: 110GB resident, 200tps prefill, only 12-15tps generation, 65k context — but output quality is "night and day" better.

Key findings:

Related event: Enthusiasts Test Local Qwen Flash Next: Memory Buys Quality but Speed Lags(2 posts)→

Original post →

More from Infra

Infra channel →