NVIDIA says ModelExpress cuts DeepSeek-V4 Pro startup from 8 minutes to under 2
NVIDIAAI · x · 2026-07-25
NVIDIA says it cut DeepSeek-V4 Pro startup time from 8 minutes to under 2 minutes by moving model weights over the fastest path to GPU memory using GPU-to-GPU RDMA.
- The company says the result comes from NVIDIA ModelExpress (MX), a weight distribution and cache management service inside NVIDIA Dynamo.
- MX reuses kernel caches and lets inference workers fetch updated weights directly from other GPUs via NIXL, avoiding centralized broadcasts.
- NVIDIA says the same approach also speeds up inference and RL post-training by keeping weight movement off the critical path.
Related event: NVIDIA ModelExpress Reduces DeepSeek Startup Time to Under 2 Minutes(2 posts)→
More from Infra
- Rocket Lab says Neutron launch capacity is now on sale — BrettKrieger12 · 2026-07-25
- Hermes Agent adds a credential firewall for Docker sandboxes — NousResearch · 2026-07-25
- Moonshot’s open-weight Kimi K3 is nearing top U.S. models and shifting AI economics — OmarUFlorez · 2026-07-25
- Morgan Stanley sees Big Tech capex hitting $1.16T in 2027 — Beth_Kindig · 2026-07-25
- Marcus says the AI trade is under stress as CRWV and ORCL slide sharply — GaryMarcus · 2026-07-25
- CachyLLama fork cuts repeated prompt processing in long local-agent sessions — UsualResult · 2026-07-25