8x RTX 3090 Setup Serves Qwen Flash Next at 661 tok/s with 262k Context

QuixiAI · x · 2026-08-27

User QuixiAI shares a benchmark running SlimServe on 8x RTX 3090s: Qwen 3.8 Flash Next delivers 150 tok/s at concurrency 1 and up to 661.1 tok/s at c32, while sustaining 262k context length — showing that retired consumer GPUs can handle long-context inference workloads.

Original post →

More from Infra

Infra channel →