Best Inference Engine for Qwen 3.8 27B on Dual RTX 5090s

youcloudsofdoom · reddit · 2026-08-18

A Reddit user seeks recommendations for the best inference engine to run Qwen 3.8-27B on a dual RTX 5090 setup. The goal is to utilize Q8/BF16 precision instead of NVFP4 to maximize decode speed, moving beyond the limitations of current single-GPU optimized recipes.

Original post →

More from Infra

Infra channel →