Running Qwen3.8-27B EXL3 on RTX 3060 + 5060 Ti: 50 tok/s with tensor parallelism and MTP
bring_back_the_v10s · reddit · 2026-09-21
A user details running Qwen3.8-27B EXL3 (5.0bpw) on a mixed-GPU rig — RTX 5060 Ti 16GB + RTX 3060 12GB (28GB total), 32GB RAM, Ryzen 5 5600, Linux — via exllamav3/tabbyAPI.
Key setup and results:
- Native tensor parallelism with auto GPU split, 102K context, Q8 cache;
- MTP draft mode lifts throughput to 50 tok/s average (40-60 swing) vs 22 tok/s without;
- Pitfalls: manual GPU-split tweaks hang the engine; 1GB VRAM used by the desktop environment contributes to imbalance (leftover: 1.3GB on the 3060, 0.6GB on the 5060);
- He already used it to write a small Rust TUI app with ratatui.
Full config.yml included — directly reusable for anyone running 27B models on older cards.
More from Infra
- Researchers formally verify the Kubernetes control plane with a compositional CORE spec — tianyin_xu · 2026-09-21
- 99.7% cache hits: engineered DeepSeek Harness with self-hosted GLM-5.3 — burny_tech · 2026-09-21
- Solo dev open-sources 4 systems projects, asks engineers to roast them — Accomplished_Row1433 · 2026-09-21
- Andrew Chen: strong LLMs are far from running on phones, on-device AI faces bandwidth, heat and model-size hurdles — andrewchen · 2026-09-21
- Tobi Lütke: local Dell server runs DeepSeek 4.1 Flash at ~300 tok/s, a billion tokens a month — BLUECOW009 · 2026-09-21
- Baseten CEO says token volume grew 40x YoY while revenue grew ~10x in 12 months — rohanpaul_ai · 2026-09-21