Running DeepSeek-V4-Flash on GB10: Setup Guide for 23+ tok/s
lilian_moraru · reddit · 2026-08-05
A developer shared a detailed guide for deploying DeepSeek-V4-Flash on the Spark GB10 edge device using llama.cpp.
- Performance: By upgrading to Nvidia driver 610.43.02 and CUDA 13.3.1, combined with UD-IQ3XSS quantization and DSpark BF16 speculative decoding, generation speeds reach 23-27 tok/s, and 25-33 tok/s for code generation.
- Setup: The post provides complete commands from firmware updates and environment variables to compiling llama.cpp from source.
- Trade-offs: To fit the 400K+ bf16 KV cache and DSpark within memory limits, the UD-IQ3XXS quantization strategy was necessary, trading some model quality for codegen speed. Prompt processing speed still has room for improvement.
More from Infra
- Anthropic Officially Confirms In-House Custom AI Chip Team for Claude — ns123abc · 2026-08-05
- Proxmox VE Launches Official ARM64 Support, Fully Compatible with NVIDIA Grace Hopper — jedisct1 · 2026-08-05
- Benchmarking MiniMax H3: Doubling Video Length Nearly Triples Render Time — jtreminio · 2026-08-05
- TensorSharp MoE Offload Slashes VRAM Use, Outperforms llama.cpp by up to 8x — fuzhongkai · 2026-08-05
- Oxford Economist: AI Arms Race Drives Massive Data Center Rollout — carlbfrey · 2026-08-05
- Cerebras Founder Breaks Down AI Chip Supply Chain and Inference Migration — mattturck · 2026-08-05