Single 3090 + 32GB RAM: squeezing max fidelity out of local Qwen3.8-27B
Certain_Yam_5824 · reddit · 2026-09-12
The OP details a local setup with a single 3090 (24GB VRAM) and 32GB DDR4, running Qwen3.8-27B at q80 with a 128k context window. To fit context in VRAM they had to strip the mmproj (multimodal) component from the model. They want a bigger overnight model but can't find multimodal options that fit 48GB after context overhead, currently rerunning Qwen3.8 at 256k when needed. The post includes concrete VRAM tradeoffs and asks for 48GB model recommendations.
More from Infra
- Dynamic llama.cpp Config Manager Pushes 27B Model From 167k to 262k Context on One 32GB GPU — wadeAlexC · 2026-09-12
- DigitalOcean Launches M.A.R.S. Managed Agent Runtime With First-Party OpenAI Agents API Support — OpenAIDevs · 2026-09-12
- Instinct may burn $100M+ a year in tokens, and open-weight models aren't actually cheaper — ivan_bezdomny · 2026-09-12
- Curie: a from-scratch 17B model designed to run from SSD, 33 tokens/s on one CPU core — Just_Vugg_PolyMCP · 2026-09-12
- LMStudio now accepts llama.cpp overrides; --yarn-attn-factor 1.2 may boost creativity — Extraaltodeus · 2026-09-12
- Analyst details Apple's S11, A20 and M6 silicon: new packaging, cooling and on-device AI designs — BenBajarin · 2026-09-12