CoreWeave RL Rollouts Hot-Loads Policy Weights Into Live Deployments, ~15x Faster Than Redeploys
_ScottCondron · x · 2026-10-07
CoreWeave launched RL Rollouts (preview) to close the inference-training loop in RL post-training: instead of redeploying per checkpoint, it transfers weight deltas and hot-loads them into a live deployment without restarting inference or touching in-flight requests — about 15x faster than a redeploy cycle.
- Built on NVIDIA Dynamo with vLLM (dynamo-vllm engine), exposed via the Dedicated Inference API.
- NVIDIA and You.com used it to post-train Nemotron 3.5 Lightning for web search, lifting BrowseComp accuracy from 36.97% to 45.45% while cutting tool calls by 30.24%.
- Announced alongside Model Distillation and a programmable training API, forming a pipeline from production signals back into model improvement.
Related event: CoreWeave Launches RL Rollouts with 15x Faster Weight Hot-Loading(2 posts)→
More from Infra
- Qualcomm denies Huawei cross-license covers LogicFolding chip patents — teortaxesTex · 2026-10-07
- Ai2 publishes technical report on supercharging Olmo-core for scalable MoE training — StasBekman · 2026-10-07
- OpenSSH drops sshd sandboxing on macOS as SDK 27 removes key API — jedisct1 · 2026-10-07
- Nvidia, Broadcom, AMD and Micron now worth a combined $9.8T, up from $1.7T three years ago — Beth_Kindig · 2026-10-07
- Cloud Run's hidden 60% scaling dials go GA: CPU and concurrency targets now customizable — rseroter · 2026-10-07
- MIT Lincoln Lab tracks 120+ AI accelerators from GPUs to ASICs in LAICS survey — nordicinst · 2026-10-07