Maximizing throughput: running parallel LLM instances on 2x V100s
Kike328 · reddit · 2026-08-30
A user running DeepSeek Q8 on dual V100s (64GB VRAM) achieves 20 tok/s. They aim to run multiple independent agents/experiments in parallel to increase aggregate throughput, asking if this is feasible given the hardware constraints and how to configure it using llama.cpp or a harness.
More from Infra
- llama.cpp NUMA mirroring boosts dual-EPYC inference by up to 137% — mattescala · 2026-08-30
- Autonomous Launches Personal AI Datacenter Hardware Starting at $26,100 — dee_hw · 2026-08-30
- Seeking the current best LLM inference setup for dual A100 GPUs — Theio666 · 2026-08-30
- Analysis: Meta's AI Infrastructure is Mispriced and Massive — RihardJarc · 2026-08-30
- FlashAccel: Leveraging High-Bandwidth Flash for LLM Inference — 9r4n4y · 2026-08-30
- Qwen3.8-27B on RTX 5090: NVFP4 Quantization Achieves 256 t/s Code Gen with 175k Context — pennyonaire · 2026-08-30