Single 3090 + 32GB RAM: squeezing max fidelity out of local Qwen3.8-27B

Certain_Yam_5824 · reddit · 2026-09-12

The OP details a local setup with a single 3090 (24GB VRAM) and 32GB DDR4, running Qwen3.8-27B at q80 with a 128k context window. To fit context in VRAM they had to strip the mmproj (multimodal) component from the model. They want a bigger overnight model but can't find multimodal options that fit 48GB after context overhead, currently rerunning Qwen3.8 at 256k when needed. The post includes concrete VRAM tradeoffs and asks for 48GB model recommendations.

Original post →

More from Infra

Infra channel →