Qwen3.8-27B Runs at 130 t/s Locally on Five-Year-Old Gaming GPUs
IgorCarron · x · 2026-08-18
AI researcher Igor Carron reshares Antoine Tilloy's hands-on test: Alibaba's Qwen3.8-27B is roughly 1% the size of frontier models, yet runs locally at 130+ tokens/s on a pair of five-year-old consumer gaming GPUs. Per Artificial Analysis benchmarks, it scores close to the frontier — though the original thread flags caveats worth noting.
More from Infra
- GitHub outages spark renewed interest in CPU-based infrastructure reliability — tokenbender · 2026-08-18
- Building a Hybrid Local-Cloud Workflow with Hermes Agents — max_paperclips · 2026-08-18
- Broadcom may surpass NVIDIA in HBM demand by 2028 — zephyr_z9 · 2026-08-18
- Minisforum's new NAS packs Strix Halo with 128GB at 8533MT/s for $3,599 — fallingdowndizzyvr · 2026-08-18
- Cherokee Nation Bans Hyperscale Data Centers on Its Lands — KeanuRave100 · 2026-08-18
- Best Inference Engine for Qwen 3.8 27B on Dual RTX 5090s — youcloudsofdoom · 2026-08-18