Local Multi-GPU Cluster Concurrently Runs Multiple Open-Source LLMs with Performance Stats
Any-Lingonberry7411 · reddit · 2026-08-07
A developer showcased a high-end local multi-GPU cluster used for agentic coding, sharing performance metrics for concurrently running multiple open-source LLMs.
- Hardware: The main rig features an RTX 6000 Pro Blackwell (96GB), 2x RTX 5090, 1x RTX 4090, and 3x AMD R9700, alongside 96GB RAM. A Strix Halo laptop and an RTX 3090 rig act as secondary machines.
- Concurrency: The setup can simultaneously run models like DeepSeek-V4-Flash (45 t/s), GLM-4.7-flash (20 t/s), and Gemma-4-12B (25 t/s). Pooling all resources allows running massive models like GLM-5.2 or Kimi-K2.7-Code.
- Bottleneck: The author notes that seamlessly switching between a "fast worker" and a "deep thinker" is currently difficult, as evicting and loading models kills the cache and slows down the workflow. They are seeking upgrade advice.
More from coding & agent
- OpenAI Dev Shares Advanced Codex Workflows for Automating Daily Tasks — petergyang · 2026-08-07
- Tencent's RST: recursive synthesis of terminal tasks at $0.05 each, boosts Qwen3.5 — Tencent-Hunyuan · 2026-08-07
- ScaleAI's HarnessOpt-Bench evaluates LLMs at optimizing agent harnesses — ScaleAI · 2026-08-07
- Testing AI Coding Agents: Bloated Instruction Files Waste 83% More Tokens — RunAI_Coder · 2026-08-07
- UltraContext: Open Source Real-Time Context Infrastructure for AI Agents — tom_doerr · 2026-08-07
- Corigin Launches ChatGPT Mobile SSH Gateway with Persistent Sandbox — mattrickard · 2026-08-07