Running Qwen Flash Next on a $1k RAM-rich GPU-poor box: 12tps but far better output
Positive-Stock6444 · reddit · 2026-08-30
A Reddit user shares benchmarks from running Qwen models locally on a 2018 ThinkStation P520 (Xeon W-2145, 256GB quad-channel DDR4, 12GB 3060, €1,000 total).
Qwen3.6 35B A3B Q4KM (former daily driver): 20GB RAM resident, 400tps prefill, 30-50tps generation, 128k context, and tolerant of other load on the box.
Qwen3.8 Flash Next UD-Q4KXL: 110GB resident, 200tps prefill, only 12-15tps generation, 65k context — but output quality is "night and day" better.
Key findings:
- Flash Next's performance collapses to 3-5tps under any memory contention; it's only really usable when opencode runs from another machine
- Synthetic random-content benchmarks gave completely wrong performance readings — realistic contexts are essential for measuring MoE models
- Full llama-server invocations for both models are included, covering GPU layer counts, CPU MoE allocation, quantized KV cache, MTP speculative decoding, and sampling params; Flash Next runs with unbounded thinking for full quality
Related event: Enthusiasts Test Local Qwen Flash Next: Memory Buys Quality but Speed Lags(2 posts)→
More from Infra
- Qwen3.8 Flash hits 415 tok/s on dual DGX Sparks — NVIDIAAI · 2026-09-01
- OpenAI's 'Jalapeno' Chip Revealed: 1500 Tokens/s Throughput — firstadopter · 2026-09-01
- TensorSharp vs llama.cpp: Qwen 3.8 Flash Next Benchmarks — fuzhongkai · 2026-09-01
- Tencent Hunyuan AngelSlim: Compressing Hy4 Model to 214GB with Heterogeneous Inference — 腾讯混元 · 2026-09-01
- Samsung shifts to 8-layer HBM4E for Nvidia with ~20% higher speed spec — 创业邦 · 2026-09-01
- Why did increasing context size increase speed in Llama.cpp? — satnl · 2026-09-01