NVIDIA and Cedana serve a portfolio of coding models on one 8x B200 node
josh_wills · x · 2026-10-02
NVIDIA and Cedana published a joint solution for hosting a full portfolio of frontier coding models on a single 8x B200 node.
- Architecture: NVIDIA Dynamo serves, NeMo Switchyard routes and escalates when a smaller model fails, and Cedana switches models in about a minute with sessions intact.
- Motivation: Enterprises self-host coding models for code privacy and cost control, but coding work spans autocomplete, small edits, complex debugging, and repo-wide planning — no single model is optimal across all.
- Key data: Per TraceLab (4,300 public Claude Code/Codex sessions), the median request finishes in 38 seconds, while 6% of requests over 10 minutes consume 69% of total request time — so a resident mid-size model serves most traffic, with the frontier model only for the tail.
- Portfolio: DeepSeek-V4-Pro-NVFP4 (873 GiB) for long-horizon planning, Kimi-K2.6 / GLM-5.26 for frontier tasks, plus smaller tiers for medium complexity.
More from coding & agent
- What should make you walk away from a voice-agent pilot? False confirmations — Legitimate-Pride-685 · 2026-10-02
- Users: new Gemini shines on open-ended overnight research tasks, Flash covers daytime — JMateosGarcia · 2026-10-02
- Claude Code Desktop Lets You Pop Diffs and Terminal Panes Out to a Second Screen — EricBuess · 2026-10-02
- AI agents play The Traitors UK in parallel: a 30-day emergent-behavior sim — bobo-the-merciful · 2026-10-02
- Two New Papers Use Coding Agents for 3D Scene Reconstruction; Gemini 3.8 Impresses Early — kwangmoo_yi · 2026-10-02
- Anthropic's Thariq Shihipar on Claude Mods, Mutable Software, and Multiplayer Agents — daniel_mac8 · 2026-10-02