Running Qwen 27B and Gemma 31B locally on one RTX 4090: quantization and context tradeoffs
MooseEfficient2151 · reddit · 2026-09-08
A Reddit user with an RTX 4090 (24GB), 64GB DDR5 and a Ryzen 7950x wants to shift code indexing and internal doc processing from external APIs to local models for privacy. Qwen 3.6 27B runs smoothly at Q4 in VRAM, but larger dense models and gpt-oss 20b with context offloading tank generation speed. They ask for vLLM/llama.cpp backend and quantization setups that balance context size and tok/s on a single card.
More from coding & agent
- Open-source AI skill clears 135GB of Mac cache overnight, saving a ¥1500 SSD upgrade — oran_ge · 2026-09-08
- Remotion ships whisper-webgpu: in-browser on-device transcription straight to captions — nicodotdev · 2026-09-08
- Stress-testing Instinct agent: snacks, flight alerts work, payments remain a trust blocker — Sad-Locksmith5980 · 2026-09-08
- How building an App Store Connect CLI in public and agents shaped a career — rudrank · 2026-09-08
- Analook MCP Connector Adds Competitor Intelligence for AI Agents — modelcontextprotocol · 2026-09-08
- YouTube-to-MP3 MCP Server Released With Custom Quality and Time-Range Conversion — modelcontextprotocol · 2026-09-08