Sizing an on-prem multi-agent coding stack: model tiers vs VRAM, and which harness

litLikeBic177 · reddit · 2026-10-07

A developer planning a fully on-prem multi-agent coding setup (H200-class GPUs, 2-3 users, no code to external APIs) asks which open-weight model tier actually matters for repo-level agentic work — 30B, 120B, or 500GB+ — citing Vibe Code Bench results showing small models collapse on long E2E builds. They also weigh heterogeneous planner/executor setups and harnesses (OpenHands, OpenCode, Cline/Roo, Codex CLI local mode) for routing sub-agents across endpoints.

Original post →

More from coding & agent

coding & agent channel →