DeepSeek V4 Flash Prefill Speed Boost: Downgrade to CUDA 13.1
fragment_me · reddit · 2026-08-03
The author shares two effective solutions to boost DeepSeek V4 Flash prefill/pp performance by 100-150 t/s.
- Option 1 (Preferred): Downgrade CUDA from 13.3 to 13.1. Starting from version 13.2, DeviceTopK is used instead of argsort for top-k operations, which ruins the PP rate.
- Option 2: Use a modified fork compatible with CUDA 13.3.
The author notes that DS4 Flash currently spends a significant amount of time on operations other than matrix multiplication.
More from Infra
- 30B Video MoE Quantizations Tested: Most Users Should Wait — EntireBig7258 · 2026-08-03
- AI Data Centers Consume Up to 1.5 Billion Gallons of Water Yearly — AndyMasley · 2026-08-03
- Prepping for Local LLM Inference: Enthusiast Builds 30TB SSD & 256GB RAM Rig — reto-wyss · 2026-08-03
- CPO Packaging Tech Unlikely to See High-Volume Shipments Before 2028 — BenBajarin · 2026-08-03
- Cornell Releases Roadmap for Parallel Programming and HPC Concepts — thehiphopswami · 2026-08-03
- Running DeepSeek V4 Flash 155G on DGX Spark: 2-bit Quantization & MTP Benchmarks — Puzzleheaded_Base302 · 2026-08-03