DIY CPU/NVMe/GPU hybrid boosts 460GB DeepSeek inference from 216 to 960 tok/s prefill

HankYeomans · x · 2026-09-25

A developer shares hands-on results running DeepSeek's MXFP4 quantized model (460GB) on a custom CPU/NVMe/GPU hybrid inference setup: prefill jumped from 216 to 960 tok/s and decode from 13 to 42 tok/s. The bottleneck is now the CPU/memory/NVMe side — only one GPU hits 100% utilization while the rest idle at 20-30%. The author argues you learn most by watching your own systems fail, and plans to next tackle end-to-end model training, optimization, and ablation from a base checkpoint.

Related event: Developer Runs 460GB DeepSeek on CPU/NVMe/GPU Hybrid, Boosting Prefill to 960 tok/s(3 posts)→

Original post →

More from Infra

Infra channel →