Running Qwen3.8 locally on a 128GB laptop for agentic coding: thinking tokens, not tok/s, set the wall clock

deepu105 · reddit · 2026-09-11

The author runs Qwen 3.8 (27B and Flash Next) on an ASUS ROG Flow Z13 (Ryzen AI Max+ 395, 128GB unified memory, Arch Linux) with llama.cpp, his own LlamaStash launcher, and Pi as the coding harness — and argues it can replace Claude Opus 4.6–4.8 for agentic coding if you accept 2–3x longer tasks.

$0/month, fully offline. Full benchmarks, configs and tuning writeup on deepu.tech.

Original post →

More from coding & agent

coding & agent channel →