Developer Runs 460GB DeepSeek on CPU/NVMe/GPU Hybrid, Boosting Prefill to 960 tok/s
A developer shared a week-long experiment running the 460GB MXFP4-quantized DeepSeek v4.1 on a self-built CPU/NVMe/GPU hybrid system, boosting prefill throughput from 216 to 960 tok/s—roughly 4.4x faster.
2026-09-25 ~ 2026-09-25 · 3 related posts
- Running DeepSeek v4.1 MXFP4 on CPU/NVMe/GPU hybrid: prefill boosted to 960 tok/s — HankYeomans · 2026-09-25
- Hybrid CPU/NVMe/GPU Setup Runs DeepSeek's 460GB MXFP4 Weights at 960 tok/s Prefill — HankYeomans · 2026-09-25
- DIY CPU/NVMe/GPU hybrid boosts 460GB DeepSeek inference from 216 to 960 tok/s prefill — HankYeomans · 2026-09-25