Running Qwen 3.8 27B on Dual RTX 3090s: 524k Context at 60-88 tok/s

elsung · reddit · 2026-09-05

Reddit user elsung shares a working local setup for Qwen 3.8 27B (HuiHui abliterated config) on dual RTX 3090s: 524k token context at c=1, roughly 60-88 tok/s, with claimed decent accuracy and Chain-of-Draft to curb overthinking. After fixing errors from an earlier post (confusion with Qwen 3.8 Flash Next), the author open-sourced the benchmark configs on GitHub for reproduction.

Original post →

More from Infra

Infra channel →