Qwen 3.8 27B Local Coding: 73 tok/s at 128K Context on RTX 4090

julianharris · x · 2026-08-17

A developer tests Qwen 3.8 27B locally on RTX 4090 with 4-bit quant and MTP, achieving 73 tok/s at 128K context. Disabling thinking gives 78 tok/s but lower quality. Compared to Qwen 3.6, setup is easier and more stable.

Related event: Qwen 3.8 27B Shows Stable Local Performance(2 posts)→

Original post →

More from coding & agent

coding & agent channel →