Can llama-server do round-robin load balancing across 3 AMD MI50s?
HlddenDreck · reddit · 2026-09-01
A user running llama-server locally on three AMD MI50s found tensor parallelism too slow due to PCIe 3.0 bandwidth and cards on different NUMA nodes. They want each GPU running the same model with server-side round-robin request distribution (request 1 to card 1, request 2 to card 2), without client-side multi-port load balancing. They ask whether llama-server alone can do this or if middleware is needed.
More from Infra
- Google details its full AI stack for developers — fhinkel · 2026-09-02
- Token Reduction Tools Don't Save Money: Cutting 38% Output Saves Only 1.3% of Bill — JFPuget · 2026-09-02
- Inference optimization startup Wafer AI raises $40M Series A — ycombinator · 2026-09-02
- David Manheim: Cost Analysis of the Hugging Face Swarm Attack — davidmanheim · 2026-09-02
- Identical ComfyUI workflow on RX 9070 XT suddenly 3-5x slower, suspected ROCm VRAM eviction — bosox62 · 2026-09-02
- Token doubling is AI's Moore's law; interactive clock puts AI at 1% of US GDP by 2028 — robleclerc · 2026-09-02