Serving a Full Coding Model Portfolio on One 8x B200 Node With ~1-Minute Switching
josh_wills · x · 2026-10-02
Cedana, working with NVIDIA's Dynamo team, shows how to host a portfolio of frontier coding models on a single 8x B200 node, switching between them in about a minute with in-flight sessions intact — solving the 9-34 minute cold-start problem that makes model portfolios impractical.
Key points:
- Enterprises increasingly run model portfolios: cheap small models for easy tasks, frontier models for the hard tail. The bottleneck is GPU utilization — models don't fit on one node together, and cold starts kill switching.
- Based on TraceLab, a public trace of 4,300 Claude Code and Codex sessions: median requests finish in 38 seconds, but the 6% running past 10 minutes consume 69% of all request time — heavily long-tailed traffic.
- The stack: NVIDIA Dynamo serves, NeMo Switchyard routes and escalates when smaller models fail, Cedana handles 1-minute model switching. Coding agents connect to a single API endpoint.
- The article lists the portfolio and VRAM footprints, e.g. DeepSeek-V4-Pro-NVFP4 at 873 GiB for long-horizon planning.
Related event: Cedana and NVIDIA host full coding model fleet on one node(2 posts)→
More from coding & agent
- Claude Given 12 Hours Autonomously Produced The Clodyssey — FinanceYF5 · 2026-10-02
- Hebbrix Launches Hosted MCP Memory Endpoint, Warns of Write-to-Searchable Gap — Separate_Sand8265 · 2026-10-02
- Agent hits first roadblock: safety filter forces manual send, says Greg Mushen — gregmushen · 2026-10-02
- HumanLayer teases 'software forge for AI' as Zed CEO jokes about the rebrand — zeeg · 2026-10-02
- Karpathy: From controlled-language docs to bespoke explainer videos, better ways to read LLM output — karpathy · 2026-10-02
- Magnetic bento hover effect using pure CSS anchor positioning — jh3yy · 2026-10-02