DeepSeek-V4-Flash hits 44-59.5 tok/s on RTX PRO 6000 eGPU with llama.cpp

backslashHH · reddit · 2026-08-02

Reddit user backslashHH shares benchmark results for running DeepSeek-V4-Flash-0731 on a Bosgame M5 mini PC with an RTX PRO 6000 Max-Q eGPU. Using llama.cpp with DSpark draft model, decode speeds are 44.0 t/s for UD-Q8KXL, 48.4 t/s for UD-Q4KXL, and 59.5 t/s for UD-Q2KXL (without drafter). Prefill speeds are 564, 585, and 1513 t/s respectively. The post includes detailed launch commands and layer split configurations.

Original post →

More from Infra

Infra channel →