NVIDIA says ModelExpress moves DeepSeek-V4 Pro startup below 2 minutes
NVIDIAAI · x · 2026-07-25
NVIDIA says its ModelExpress (MX) service inside Dynamo can cut DeepSeek-V4 Pro startup time from 8 minutes to under 2 minutes.
- MX moves weights over GPU-to-GPU RDMA and reuses kernel caches.
- It avoids centralized broadcasts by letting inference workers fetch updated weights directly from other GPUs via NIXL.
- NVIDIA says the same design speeds up both inference and RL post-training by keeping weight transfer off the critical path.
Related event: NVIDIA ModelExpress Reduces DeepSeek Startup Time to Under 2 Minutes(2 posts)→
More from Infra
- GitHub Pull Requests hits an outage with degraded availability and create errors — 0xkarasy · 2026-07-25
- Schmidhuber marks Nvidia’s $4 trillion milestone and says compute is 100,000× cheaper — SchmidhuberAI · 2026-07-25
- Rocket Lab says Neutron launch capacity is now on sale — BrettKrieger12 · 2026-07-25
- NVIDIA says ModelExpress cuts DeepSeek-V4 Pro startup from 8 minutes to under 2 — NVIDIAAI · 2026-07-25
- Hermes Agent adds a credential firewall for Docker sandboxes — NousResearch · 2026-07-25
- Moonshot’s open-weight Kimi K3 is nearing top U.S. models and shifting AI economics — OmarUFlorez · 2026-07-25