SlimServe runs Qwen 3.8 Flash Next on 8x RTX 3090s at up to 1200 tok/sec decode
QuixiAI · x · 2026-09-06
The open-source inference project SlimServe (built on ds4/vllm) benchmarks Qwen 3.8 Flash Next on 8x RTX 3090s with P2P enabled: 140 tok/sec decode at C1 and 1200 tok/sec at C32. The author calls it "an incredible amount of valuable inference on modest hardware," offering a low-cost local-serving reference.
More from Infra
- Musk: 3D printing enables integrated flow paths but is too slow and costly for volume production — i_bioloid · 2026-09-06
- Nvidia guides 70% revenue growth next year, supply sets the ceiling — BenBajarin · 2026-09-06
- Unreal Engine pipeline yields 8.7K+ hours of action-conditioned video for world-model pretraining — udmrzn · 2026-09-06
- PyPI's Recent Download Corrections Sharply Cut Some Packages' Stats — dbreunig · 2026-09-06
- Running gemma4:31b-mlx locally with Ollama feels indistinguishable from paid tiers — walkingriver · 2026-09-06
- Redditor seeks a lightweight OpenAI-compatible API client, complains OpenWebUI is 30GB — MelodicRecognition7 · 2026-09-06