Dual R9700 vLLM Setup Halves Speed with Multiple Instances
Certain_Series6810 · reddit · 2026-08-23
A user compared llama.cpp and vLLM on a dual R9700 Ubuntu setup. While llama.cpp achieved 40 t/s for a single instance, switching to the vLLM Docker image to support multiple instances halved the generation speed. The user asks if this is expected behavior for vLLM.
More from Infra
- Run DeepSeek-V3 on $47 Hardware with Pruning and Quantization — StefanoGogioso · 2026-08-23
- Call for Top Architects to Design Beautiful Datacenter Exteriors — EddyVGG · 2026-08-23
- 25% chance orbital data centers launch by end of next year — Polymarket · 2026-08-23
- VC proposes network of giant data centers along US-Mexico border — Polymarket · 2026-08-23
- Experiment proposed: Local Qwen model on Mac vs $10k cloud security scan — natesiggard · 2026-08-23
- Best Local LLMs for 12GB VRAM: Alternatives to GLM 4.7 Flash? — OrangeThink5911 · 2026-08-23