DeepSeek-V3-Flash Runs at 47 tok/s on a Single $47 Setup
A GitHub project demonstrates running DeepSeek-V3-Flash on a single DGX Spark device with pruning and EXL3 quantization, achieving 47 tok/s generation and 400K context for about $47.
2026-08-21 ~ 2026-08-23 · 2 related posts
- DeepSeek v4 runs at 47 tok/s on single GPU, passes 370k needle test — EAccelerate_42 · 2026-08-21
- Run DeepSeek-V3 on $47 Hardware with Pruning and Quantization — StefanoGogioso · 2026-08-23