Hybrid CPU/NVMe/GPU Setup Runs DeepSeek's 460GB MXFP4 Weights at 960 tok/s Prefill

HankYeomans · x · 2026-09-25

The author shares a week of local inference experiments: running DeepSeek v4.1's MXFP4 quantized weights (460GB) on a self-built CPU/NVMe/GPU hybrid, boosting prefill from 216 tok/s to 960 tok/s and decode from 13 tok/s to 42 tok/s.

Key finding: the current bottleneck isn't the GPUs but the CPU/memory/NVMe side — only 1 of 20 GPUs reaches 100% utilization while the rest sit at 20-30%. Next, the author plans end-to-end model creation, optimization and ablation from a base checkpoint, arguing the only way to learn is to do.

Related event: Developer Runs 460GB DeepSeek on CPU/NVMe/GPU Hybrid, Boosting Prefill to 960 tok/s(3 posts)→

Original post →

More from Infra

Infra channel →