Optimizing Qwen3.8-27B: From 9.5 to 153 Tokens/Second

TrifleHopeful5418 · reddit · 2026-08-21

Through 159 experiments, the author optimized Qwen3.8-27B on a heterogeneous setup (AMD Strix Halo + RTX 3090 Ti), boosting generation speed from 9.5 to 153 tok/s (32K code context) and outperforming a dual-3090 vLLM cluster on HumanEval.

Key Optimizations:

Failed Approaches:

Benchmarks:

Original post →

More from Infra

Infra channel →