Running 27B Models on a Single RTX 5090: Multi-GPU vs. Edge Deployment Stats
Fz1zz · reddit · 2026-08-13
After acquiring an RTX 5090, a developer shared performance metrics and hardware advice for local LLM deployment.
- Main Rig: Running the Qwen3.6-27B model on a single 5090 with a 160k context, Flash Attention, and Multi-Token Prediction (MTP). The author argues that adding a 4070 Ti Super for a multi-GPU setup would likely slow things down rather than help, preferring to sell the old card.
- Edge Device: On a NucBox K8 Plus mini PC with an AMD Ryzen 7 8845HS, utilizing the iGPU and unified memory to run Gemma-4-26B, the setup achieved 343 t/s prefill and 31.4 t/s decode speeds.
More from Infra
- Red Hat's DSpark Speculator Boosts Kimi-K3 Throughput by 3.5x — teortaxesTex · 2026-08-14
- AMD Raising $5 Billion in Debt to Fund AI War Against Nvidia — ns123abc · 2026-08-14
- Merge Gateway Launches Multimodal API for Unified Access to Image, Video, Audio Models — shensi · 2026-08-14
- MasterClass Adopts CoreWeave and W&B Weave to Monitor AI Teaching Agents — wandb · 2026-08-14
- SkyPilot AI Infra Meetup next Tuesday in SF with VAST Data and NVIDIA — skypilot_org · 2026-08-14
- SanDisk Predicts Flash Market to Approach $500B by 2027 Amid AI Boom — zephyr_z9 · 2026-08-13