Running a 27B LLM on three RTX 5060 Ti cards: one user's local MoA setup diary
Then_Blueberry7290 · reddit · 2026-09-05
A Reddit user details a full local AI deployment: a Dell 5820 workstation (Xeon w2245) with two RTX 5060 Ti 16GB cards runs NVFP4-quantized Qwen3.8-27B with 210k context, Q8 KV cache, and vision in VRAM at 38-44 tokens/s.
Adding a third card required an external riser, degraded cooling and stability, and actually slowed prompt processing under full tensor parallelism. The user is now exploring a second GPU-less PC to host the third card, testing Ornith 1.5 35B (17-20 t/s) and Gemma 4 26B NVFP (30-40 t/s) at 16GB VRAM.
Production use cases include WordPress site-maintenance agents, office document handling, and generating prompts for image/video models like Z-Image Turbo and LTX 2.5 driven through ComfyUI. Current Mixture of Agents: Gemma 4 26B as reference, Qwen 3.8 27B as aggregator — with the caveat that Gemma models are lazy at tool calling.
More from coding & agent
- Squad adds GPT-6 Astra same-day, pitches model-agnostic AI teammates across 11 providers — tibo_maker · 2026-09-05
- Data scientist: 90% of peers lack confidence in time series forecasting — a 5-concept primer — mdancho84 · 2026-09-05
- Hugging Face details Moon Bot, its Slack-native coding agent with codebase and DB access — victormustar · 2026-09-05
- Indie dev expands Australian business-day MCP into 4-tool date engine, now on official MCP Registry — Impossible_Bit_2676 · 2026-09-05
- Computer use is now where coding agents were in Nov 2025, says agentic tooling author — austinvhuang · 2026-09-05
- Dev boosts GLM 5.2 TPS on a B300 and swaps it into Claude Code in place of Anthropic models — abhijithneil · 2026-09-05