8x RTX 3090 Setup Serves Qwen Flash Next at 661 tok/s with 262k Context
QuixiAI · x · 2026-08-27
User QuixiAI shares a benchmark running SlimServe on 8x RTX 3090s: Qwen 3.8 Flash Next delivers 150 tok/s at concurrency 1 and up to 661.1 tok/s at c32, while sustaining 262k context length — showing that retired consumer GPUs can handle long-context inference workloads.
More from Infra
- ik_llama.cpp monthly update: DSpark speculative decoding, Vulkan IQ4 support, and more — pmttyji · 2026-08-27
- Test: OX Alpha runs on WebGPU for just $0.016 — yuwen_lu_ · 2026-08-27
- Local AI trade-off: 96GB mixed RAM vs. speed — QuirksNFeatures · 2026-08-27
- PyTorch Ecosystem Adds Perforated, TokenSpeed, and 8 Others — zhyncs42 · 2026-08-27
- Jalapeño chip delivers up to 1.9x efficiency, docs criticized — Artistic_Phone9367 · 2026-08-27
- OpenAI Internal Compromise Deemed More Critical than Hugging Face Incident — sjgadler · 2026-08-27