A $20 Ethernet link is enough for 39.7GB multi-node GPU inference
Chuyito · reddit · 2026-07-23
A Reddit post argues that multi-node GPU inference does not require expensive networking: a simple point-to-point Ethernet setup was enough to run a 39.7 GB laguna Q2KXL model on 2x4060 + 1x4060.
The author shares timing data across different ubatch-size settings and concludes that:
- point-to-point networking keeps traffic between the two nodes instead of the switch
- device=rpc0/rpc1 can limit worker CPU usage if you do not want the host CPU involved
- ubatch=768 was the sweet spot for this setup, trading prompt-processing speed against generation throughput
More from coding & agent
- Claude turns team names into metadata and jokes about collision-free routing — draginol · 2026-07-23
- Claude plus Chromium and Playwright MCP can now record demo videos for you — yenkel · 2026-07-23
- Frontend work with Claude Design and Claude Code is “actually so good” — trq212 · 2026-07-23
- Multiple LLMs on one project start treating collaboration structure as a core feature — repligate · 2026-07-23
- Blocks says its agent-friendly UI system is nearing real-time on-brand generation — round · 2026-07-23
- Codex needing iTunes access becomes the latest AI agent meme — KarelDoostrlnck · 2026-07-23