Troubleshooting Multi-GPU Support for Deepseek V3 in llama.cpp

erazortt · reddit · 2026-08-22

A user is struggling to optimize llama.cpp for Deepseek V3 (120GB quantized) across two GPUs (Blackwell 5000 48GB + RTX 3090 24GB). Attempts to use --split-mode layer or manual tensor offloading resulted in allocation failures or slower speeds compared to running on a single GPU with CPU offload. The post seeks advice on correct multi-GPU configuration.

Original post →

More from coding & agent

coding & agent channel →