Splitting GPUs Across VMs Boosted Ollama Performance by ~3x
NicolaZanarini533 · reddit · 2026-08-24
The author shares an infrastructure optimization tip: splitting 3 pooled GPUs (2x RTX PRO 4000, 1x RTX PRO 2000) from a single VM into two VMs (1 dedicated + 2 pooled). Benchmark results show Qwen2.5-72B throughput jumped from 12.2 tok/s to 33.91 tok/s, and muse-glimmer-30B improved from 14.3 tok/s to 22.61 tok/s. This change challenged the assumption that pooling doesn't impact speed much, demonstrating the benefits of resource isolation under load.
More from Infra
- AMD Instinct MI210 for Local LLMs: A Cost-Effective Choice? — OvertaxedOne · 2026-08-24
- Wells Fargo sees Broadcom AI chip revenue at $205B by FY28, far above consensus — Beth_Kindig · 2026-08-24
- Disaggregated LPDDR memory: A new path for AI infrastructure — jwt0625 · 2026-08-24
- Study notes: how speculative decoding accelerates LLM inference without quality loss — helloiamleonie · 2026-08-24
- Global AI Power Demand Projected to Surge 1,100% by 2033 — KyeGomezB · 2026-08-24
- RTX 3080 Memory OC to +1200 Boosts Flux Generation Efficiency — MakionGarvinus · 2026-08-24