Qwen3.8 27B Hits 39 tok/s on RTX 4080 Super, Good for Local Coding

BopSupreme · reddit · 2026-08-17

A user runs Qwen3.8-27B on an RTX 4080 SUPER 16GB, achieving 39 tok/s fully GPU-offloaded, 200-230ms TTFT on short tasks, and 34 tok/s on a 26K-token workload. It's surprisingly capable for local coding/agent work, but quality is inconsistent, not a full replacement for hosted models.

Related event: Qwen3.8-27B Benchmarks: 672 tps on RTX 3090(6 posts)→

Original post →

More from coding & agent

coding & agent channel →