Homelab With 4x RTX 4090 Weighs vLLM+P2P Patch vs llama.cpp for Qwen Models

dowitex · reddit · 2026-09-07

A homelab user with four RTX 4090s (64GB RAM, Threadripper Pro, all PCIe x16) is choosing between two local inference setups for coding: (1) vLLM + Qwen 3.8 27B dense fp8 + 256K KV cache with a new open-source-driver P2P patch that only benefits vLLM, vs (2) llama.cpp + Qwen flash next MoE iq4xs + 8-bit 200K KV cache. vLLM lacks 4-bit quants and VRAM is too tight for fp8 flash next, so the question is whether flash next's extra intelligence outweighs vLLM's higher throughput.

Original post →

More from Infra

Infra channel →