SlimServe runs Qwen 3.8 Flash Next on 8x RTX 3090s at up to 1200 tok/sec decode

QuixiAI · x · 2026-09-06

The open-source inference project SlimServe (built on ds4/vllm) benchmarks Qwen 3.8 Flash Next on 8x RTX 3090s with P2P enabled: 140 tok/sec decode at C1 and 1200 tok/sec at C32. The author calls it "an incredible amount of valuable inference on modest hardware," offering a low-cost local-serving reference.

Original post →

More from Infra

Infra channel →