Baseten cuts delta weight syncs for frontier models to under 40 seconds
baseten · x · 2026-09-09
Baseten published an engineering article on optimizing delta weight syncs for RL managed rollouts: frontier labs publish new weights, and Baseten picks them up across a global fleet so rollouts keep generating against the latest policy. For large frontier models like GLM-5.3, a new policy goes live in under 40 seconds with only a 6-second request pause. Engineer Stefanopopoulos details how the optimization works.
More from Infra
- NVIDIA ships CUDA Python 1.0 with stable APIs, making Python first-class for CUDA — PyTorch · 2026-09-09
- Inference is turning GPU compute into a tradable commodity — ArtificialAnlys · 2026-09-09
- Cohere open-sources megakernel serving engine, up to 1.58x faster than vLLM — cohere · 2026-09-09
- Solving Navier-Stokes cost 130B output tokens — up to $18M depending on model pricing — mitsuhiko · 2026-09-09
- Dell says DRAM, NAND shortages persist and nearly all leading-node products are constrained — Beth_Kindig · 2026-09-09
- Alphabet's CapitalG backs AI chip startup Celero at $3 billion valuation — dinabass · 2026-09-09