Same GPUs, 100x Gap: vLLM Hits 110 tok/s on Dual RTX 5090s Where llama.cpp Took 30 Minutes
vllm_project · x · 2026-10-07
- The vLLM team highlighted a long-context serving benchmark from @NicWAI: on dual RTX 5090s, vLLM sustained 110 tokens/s, while llama.cpp on a single 5090 took 30 minutes to first token on the same model.
- Same silicon, roughly 100x difference — long-context serving stresses memory and execution paths, so serving engine architecture matters as much as raw hardware.
- Takeaway: benchmark your inference software before buying GPUs.
More from Infra
- Nvidia nears $6T market cap with record $150B buyback; SpaceX in talks to borrow $40B for chips — rohanpaul_ai · 2026-10-07
- 7 minutes per motor, ~$5 labor cost: why robotics automation is the only path for Western manufacturing — avlok · 2026-10-07
- Railroad buildout ran at 1.5-2% of GDP for 60 years — AI capex only hit that level this year — toptickcrypto · 2026-10-07
- Paired 4:8 sparsity gives 1.35-1.65x over dense NVFP4 on B200, 1.18x in serving — vllm_project · 2026-10-07
- kipply's July-August digest: export controls lifted, alignment-faking paper, and more — kipperrii · 2026-10-07
- Marvell Investor Day 2026 materials land, with a nudge to fix the chart arrows — jwt0625 · 2026-10-07