Qwen3.8-27B on 16GB VRAM: 50 tok/s with 85k Context, Config Included

brainExploded99 · reddit · 2026-08-19

A Reddit user shares running Qwen3.8-27B on an Nvidia 5070Ti 16GB, achieving 85k context and 50 tok/s by removing the MTP layer and tuning config. Includes full config and tips like q8 cache and ubatch-size ablation.

Original post →

More from coding & agent

coding & agent channel →