Same Qwen3.8-27B, five inference recipes: decode speed ranges 33-161 tok/s
Kyrannio · x · 2026-09-25
Running the same Qwen3.8-27B across five inference recipes on a 171-item suite, Brjen found scores all clustered at 135-137 while decode speed varied from 33 to 161 tok/s, with a 1080ti + 7800xt combo leading on quality.
Takeaway: quality is set by the model, speed is set by the recipe - quantization, KV cache, speculative decoding and engine choices matter far more than swapping models.
More from Infra
- Huawei 960 racks hit 200kW per cabinet in 2027, still 4x behind Blackwell — teortaxesTex · 2026-09-25
- 10-year Treasury yield hits pre-GFC high as AI capex borrows at nation scale — tszzl · 2026-09-25
- Supply board shows HBM, Blackwell, NAND and transformers all near critical — tengyanAI · 2026-09-25
- Enterprise SSD lead times hit 16 weeks as Seagate books nearline capacity into 2028 — tengyanAI · 2026-09-25
- Asianometry: High Bandwidth Flash — what is it good for? — asssuber · 2026-09-25
- AVO agents evolve attention kernels beating FlashAttention-4 by 10.5% on B200 — bingxu_ · 2026-09-25