Qwen3.8-27B Runs at 130 t/s Locally on Five-Year-Old Gaming GPUs

IgorCarron · x · 2026-08-18

AI researcher Igor Carron reshares Antoine Tilloy's hands-on test: Alibaba's Qwen3.8-27B is roughly 1% the size of frontier models, yet runs locally at 130+ tokens/s on a pair of five-year-old consumer gaming GPUs. Per Artificial Analysis benchmarks, it scores close to the frontier — though the original thread flags caveats worth noting.

Original post →

More from Infra

Infra channel →