DIY CPU/NVMe/GPU hybrid boosts 460GB DeepSeek inference from 216 to 960 tok/s prefill
HankYeomans · x · 2026-09-25
A developer shares hands-on results running DeepSeek's MXFP4 quantized model (460GB) on a custom CPU/NVMe/GPU hybrid inference setup: prefill jumped from 216 to 960 tok/s and decode from 13 to 42 tok/s. The bottleneck is now the CPU/memory/NVMe side — only one GPU hits 100% utilization while the rest idle at 20-30%. The author argues you learn most by watching your own systems fail, and plans to next tackle end-to-end model training, optimization, and ablation from a base checkpoint.
More from Infra
- Perplexity Launches Fast Search API: 95% of Results in Under 230ms on Rust-Based Photon — perplexity_ai · 2026-09-25
- 100B Model Trained Across 5 Data Centers on Plain Internet Links at 30.8% MFU — markjeffrey · 2026-09-25
- AMD gaining 10 points of GPU share would be 'transformational', analyst argues — Beth_Kindig · 2026-09-25
- Google sees orbital AI data centers reaching cost parity with terrestrial ones by mid-2030s — McDonaghMatthew · 2026-09-25
- Goldman Sachs hikes AI power forecasts: 2030 data center capacity raised to 217GW — McDonaghMatthew · 2026-09-25
- Puro-2B: an open recipe trains a Qwen2-1.5B-beating LLM on RTX 5090s for just $4.4K — IgorCarron · 2026-09-25