Developer Runs 460GB DeepSeek on CPU/NVMe/GPU Hybrid, Boosting Prefill to 960 tok/s

A developer shared a week-long experiment running the 460GB MXFP4-quantized DeepSeek v4.1 on a self-built CPU/NVMe/GPU hybrid system, boosting prefill throughput from 216 to 960 tok/s—roughly 4.4x faster.

2026-09-25 ~ 2026-09-25 · 3 related posts