Draft model hits ~60 tok/s running Qwen3.8-27B at 131k context on a 16GB GPU

pneuny · reddit · 2026-09-12

Original post →

More from Infra

Infra channel →