Maximizing throughput: running parallel LLM instances on 2x V100s

Kike328 · reddit · 2026-08-30

A user running DeepSeek Q8 on dual V100s (64GB VRAM) achieves 20 tok/s. They aim to run multiple independent agents/experiments in parallel to increase aggregate throughput, asking if this is feasible given the hardware constraints and how to configure it using llama.cpp or a harness.

Original post →

More from Infra

Infra channel →