Multi-GPU Full VRAM Deployment of DeepSeek V4 Yields Only 600 t/s PP
fragment_me · reddit · 2026-08-02
A developer loaded the IQ3XXS quantized version of DeepSeek V4 entirely into VRAM across 5 consumer GPUs (2x 3080, 2x 3090, 5090), but found the Prompt Processing (PP) speed to be a sluggish 600 t/s, making it unusable as a daily driver.
The author verified the model was fully in VRAM and updated CUDA/NCCL/Llamacpp, and is now asking the community for baseline PP numbers to troubleshoot the bottleneck.
More from Infra
- Power Shortage Becomes the New Bottleneck for the AI Race Beyond Chips — ingliguori · 2026-08-02
- AMD MI355X vLLM Beats Nvidia B200 on Kimi K2.5 Inference — marksaroufim · 2026-08-02
- Together AI's Monthly Token Volume Hits 400 Trillion, Marking 10,000x Growth — togethercompute · 2026-08-02
- DeepSeek V4 Flash 3-bit Quantization Tested on 3x RTX 3090: 119GB VRAM — consultkitapp · 2026-08-02
- Is 12GB VRAM Enough? Choosing GPUs for Local LLM Deployment — rettdit · 2026-08-02
- Report: SpaceX Plans 4GW Compute Cluster, Rivaling US Grid Expansion — BenBajarin · 2026-08-02