Running Qwen3.8 27B on 3060+3080: 26.8t/s distributed setup guide
Fieser_Fettsack · reddit · 2026-08-19
A user shares a successful distributed setup running Qwen3.8-27B across RTX 3060 and 3080 via llama.cpp RPC, achieving 26.87 t/s decode speed. Includes detailed performance metrics and Docker Compose configuration.
More from Infra
- LLM Inference Engineering: From KV Cache to vLLM and SGLang — techNmak · 2026-08-19
- DFlash 2: Qwen3.8-27B hits 70 tok/s on MacBook with 4.6x speedup — songhan_mit · 2026-08-19
- NVIDIA H100 Concurrency Response of Plain Global Loads Analyzed — ssh4net · 2026-08-19
- Using HBF for KV Cache Offload Risks Endurance Burnout — zephyr_z9 · 2026-08-19
- Considered nuclear startup funded by hyperscalers, impressed by serious energy buildout — JacquesThibs · 2026-08-19
- Qwen3.8-27B on 2x 3090 hits 218 tok/s decode with vLLM + DFlash2 spec-decode — xjx546 · 2026-08-19