DeepSeek-V4-Flash hits 44-59.5 tok/s on RTX PRO 6000 eGPU with llama.cpp
backslashHH · reddit · 2026-08-02
Reddit user backslashHH shares benchmark results for running DeepSeek-V4-Flash-0731 on a Bosgame M5 mini PC with an RTX PRO 6000 Max-Q eGPU. Using llama.cpp with DSpark draft model, decode speeds are 44.0 t/s for UD-Q8KXL, 48.4 t/s for UD-Q4KXL, and 59.5 t/s for UD-Q2KXL (without drafter). Prefill speeds are 564, 585, and 1513 t/s respectively. The post includes detailed launch commands and layer split configurations.
More from Infra
- Kernel-level optimizations for DeepSeek and Laguna models on Apple Silicon — gajesh · 2026-08-02
- Apple Silicon inference runs 137% faster as MLX Challenge pushes edge limits — gajesh · 2026-08-02
- NVIDIA's Rubin GPUs Expected to Cut Inference Costs by 90% by Late 2026 — haider1 · 2026-08-02
- Choosing a GPU Cloud Provider for Production: Beyond Price — 9ds996Dev · 2026-08-02
- Rust Compiler Performance: rustdoc Build Time Slashed by 28% — charliermarsh · 2026-08-02
- Compute Still King: French AI Circle Reflects on Efficiency Gap with DeepSeek — AymericRoucher · 2026-08-02