Same Qwen3.8-27B, five inference recipes: decode speed ranges 33-161 tok/s

Kyrannio · x · 2026-09-25

Running the same Qwen3.8-27B across five inference recipes on a 171-item suite, Brjen found scores all clustered at 135-137 while decode speed varied from 33 to 161 tok/s, with a 1080ti + 7800xt combo leading on quality.

Takeaway: quality is set by the model, speed is set by the recipe - quantization, KV cache, speculative decoding and engine choices matter far more than swapping models.

Original post →

More from Infra

Infra channel →