Running DeepSeek v4.1 MXFP4 on CPU/NVMe/GPU hybrid: prefill boosted to 960 tok/s
HankYeomans · x · 2026-09-25
A developer ran the 460GB DeepSeek v4.1 MXFP4 quantized model on a CPU/NVMe/GPU hybrid setup, tuning prefill throughput from 216 tok/s to 960 tok/s and decode from 13 tok/s to 42 tok/s. He notes greater efforts exist elsewhere, but this was a hands-on learning project. Next up: end-to-end model creation, optimization and ablation from a base checkpoint.
More from Infra
- Perplexity Launches Fast Search API: 95% of Results in Under 230ms on Rust-Based Photon — perplexity_ai · 2026-09-25
- 100B Model Trained Across 5 Data Centers on Plain Internet Links at 30.8% MFU — markjeffrey · 2026-09-25
- AMD gaining 10 points of GPU share would be 'transformational', analyst argues — Beth_Kindig · 2026-09-25
- Google sees orbital AI data centers reaching cost parity with terrestrial ones by mid-2030s — McDonaghMatthew · 2026-09-25
- Goldman Sachs hikes AI power forecasts: 2030 data center capacity raised to 217GW — McDonaghMatthew · 2026-09-25
- Puro-2B: an open recipe trains a Qwen2-1.5B-beating LLM on RTX 5090s for just $4.4K — IgorCarron · 2026-09-25