Running DeepSeek v4.1 MXFP4 on CPU/NVMe/GPU hybrid: prefill boosted to 960 tok/s

HankYeomans · x · 2026-09-25

A developer ran the 460GB DeepSeek v4.1 MXFP4 quantized model on a CPU/NVMe/GPU hybrid setup, tuning prefill throughput from 216 tok/s to 960 tok/s and decode from 13 tok/s to 42 tok/s. He notes greater efforts exist elsewhere, but this was a hands-on learning project. Next up: end-to-end model creation, optimization and ablation from a base checkpoint.

Related event: Developer Runs 460GB DeepSeek on CPU/NVMe/GPU Hybrid, Boosting Prefill to 960 tok/s(3 posts)→

Original post →

More from Infra

Infra channel →