Multi-GPU Full VRAM Deployment of DeepSeek V4 Yields Only 600 t/s PP

fragment_me · reddit · 2026-08-02

A developer loaded the IQ3XXS quantized version of DeepSeek V4 entirely into VRAM across 5 consumer GPUs (2x 3080, 2x 3090, 5090), but found the Prompt Processing (PP) speed to be a sluggish 600 t/s, making it unusable as a daily driver.

The author verified the model was fully in VRAM and updated CUDA/NCCL/Llamacpp, and is now asking the community for baseline PP numbers to troubleshoot the bottleneck.

Original post →

More from Infra

Infra channel →