Four RTX Pro 6000s, 384GB VRAM—and 71% of agent time is outside the model call

BLUECOW009 · x · 2026-09-07

A developer built his biggest local rig yet—four RTX Pro 6000s with 384GB of VRAM—for running coding agents, and made a video about rethinking what the machine actually needs.

Key finding: even with four GPUs, 71% of a measured agent turn is spent outside the model call—code has to compile and pass tests. The bottleneck for local coding agents isn't inference compute but the verification loop, so the rest of the machine and environment setup matter just as much.

Related event: Dev builds 4x RTX Pro 6000 rig with 384GB VRAM for coding agents(2 posts)→

Original post →

More from coding & agent

coding & agent channel →