Dual R9700 vLLM Setup Halves Speed with Multiple Instances

Certain_Series6810 · reddit · 2026-08-23

A user compared llama.cpp and vLLM on a dual R9700 Ubuntu setup. While llama.cpp achieved 40 t/s for a single instance, switching to the vLLM Docker image to support multiple instances halved the generation speed. The user asks if this is expected behavior for vLLM.

Original post →

More from Infra

Infra channel →