Qwen3.8 27B Hits 39 tok/s on RTX 4080 Super, Good for Local Coding
BopSupreme · reddit · 2026-08-17
A user runs Qwen3.8-27B on an RTX 4080 SUPER 16GB, achieving 39 tok/s fully GPU-offloaded, 200-230ms TTFT on short tasks, and 34 tok/s on a 26K-token workload. It's surprisingly capable for local coding/agent work, but quality is inconsistent, not a full replacement for hosted models.
Related event: Qwen3.8-27B Benchmarks: 672 tps on RTX 3090(6 posts)→
More from coding & agent
- Dev asks for pricing advice on Google Apps Script label automation tool — Responsible-Box-4905 · 2026-08-17
- Leaked workflow reveals how big studios make $2M AI movies with Seedance 2.5 on Higgsfield — EXM7777 · 2026-08-17
- Teknium to update Hermes for GPT-5.6 support — Teknium · 2026-08-17
- Benchmark reveals LLMs struggle with real-world tasks despite coding prowess — 数字生命卡兹克 · 2026-08-17
- Built a full personal finance tracker using Codex in one afternoon — alexgoughcooper · 2026-08-17
- Tutorial: Building a 3D Game with Grok, Cursor, and Blender — chongdashu · 2026-08-17