Together AI adds canary rollouts for zero-downtime model upgrades on dedicated inference
togethercompute · x · 2026-09-23
Together AI launched canary rollouts on its Dedicated Model Inference offering, letting users upgrade the model behind a live endpoint with no downtime.
- Traffic shifts from the old deployment to the new checkpoint in gated steps (default 5% → 25% → 50% → 100%)
- Health checks run before each shift, and metric gates compare the new model's p95 latency and error rate against the old one
- If a gate trips, the rollout pauses at the canary share and waits for the user to resume, promote to 100%, or roll back
- Three strategies are available: canary, blue-green, and rolling, via tg CLI, REST API, and Python SDK
Related event: Together AI Launches Canary Rollouts for Zero-Downtime Model Upgrades(2 posts)→
More from Infra
- DeepSeek unveils DSec sandbox infra running 3M agent environments per day — jiqizhixin · 2026-09-23
- Qualcomm goes agent-centric: Snapdragon 8 Elite Gen 6 and agent-native devices — jiqizhixin · 2026-09-23
- Unsloth Desktop Hotfix Adds Qwen-Image-2.1 Image Editing and Fixes GGUF Loading — danielhanchen · 2026-09-23
- Qwen 3.6 35B-A3B Q6 hits ~50 tok/s on a 128GB Strix Halo — what's the best local model now? — jankeydankey · 2026-09-23
- Dedicated Hardware for Running AI Agents at Scale Arrives — cyrilzakka · 2026-09-23
- Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning — cephaloform · 2026-09-23