Squeezing Qwen 3.8 27B MTP Q8_0 + Vision into 48GB VRAM

Creative-Type9411 · reddit · 2026-08-17

A user shared specific parameters and configurations for running the Qwen 3.8 27B MTP Q80 + Vision model within 48GB VRAM (3x Tesla T4). By fine-tuning settings (e.g., disabling mmap, using mlock, draft-mtp), they achieved 19.2k context at 35t/s without offloading, asking others for optimization tips.

Original post →

More from Infra

Infra channel →