Sizing an on-prem multi-agent coding stack: model tiers vs VRAM, and which harness
litLikeBic177 · reddit · 2026-10-07
A developer planning a fully on-prem multi-agent coding setup (H200-class GPUs, 2-3 users, no code to external APIs) asks which open-weight model tier actually matters for repo-level agentic work — 30B, 120B, or 500GB+ — citing Vibe Code Bench results showing small models collapse on long E2E builds. They also weigh heterogeneous planner/executor setups and harnesses (OpenHands, OpenCode, Cline/Roo, Codex CLI local mode) for routing sub-agents across endpoints.
More from coding & agent
- raindrop_ai Cofounder Ben Hylak on Rogue AI Agents, Catching Agent Failures and What Safety Talk Misses — soleio · 2026-10-07
- IR4RL turns intermediate render progress into RL rewards, new SOTA for image-to-code — phillip_isola · 2026-10-07
- Ramp Data Shows Enterprise AI Adoption at Peak: How to Turn Your Skills into Agents — vasuman · 2026-10-07
- Combining OpenAI's Decisions API with Live API Lets Voice Agents Act Mid-Conversation — pbbakkum · 2026-10-07
- 27B Model at 256k Context, 110+ tok/s on a Single RTX 5090 via focus-llama — Ok-Shower7286 · 2026-10-07
- Google Testing Blog: Two-Way Doors — Don't Code Yourself into a Corner — rseroter · 2026-10-07