Running Qwen3.8-27B for coding on 32GB VRAM — what local LLMs do you use and why?
theexile1337 · reddit · 2026-09-08
A Redditor describes running Qwen3.8-27B (Q4/Q6 depending on context needs) for coding on a 32GB VRAM rig, plus Gemma4 31B for research and general queries to avoid sending data to cloud models. They ask whether a site exists for comparing local LLMs by use case and invite others to share their model/quantization/VRAM setups.
More from coding & agent
- Spotify cut Claude Code tokens 90% with a 350-line routing rule, not a model breakthrough — krishnan · 2026-09-08
- A developer proposes an informal agent-native mathlib, tentatively named mathgraph — Sauers_ · 2026-09-08
- Ix builds a persistent symbol graph of your codebase across 26 languages for humans and AI — tom_doerr · 2026-09-08
- Dev builds FreeBuff MCP to route ChatGPT tasks to free agent models — Swimming_Ask3859 · 2026-09-08
- GitHub now classifies agent policy blocks as 'skipped', not failures — Crescitaly · 2026-09-08
- Companion launches iMessage AI agent that reuses your ChatGPT account with MCP support — Scobleizer · 2026-09-08