Local 27B Agent on 12GB VRAM: Qwen3.8 config and long-context practice
PyaesoneP · reddit · 2026-08-24
The author shares practical experience running Qwen 3.8 27B (UDQ4KXL) for local agent coding on an RTX 5070 Ti Mobile (12GB). The setup includes 100K context, specific llama-server launch parameters, and inference settings. To handle context overflow caused by verbose reasoning, the author adopted Magic Context over native compaction, scaling sessions to 3.7M processed tokens and shipping end-to-end features. A hybrid workflow using Claude Opus for final PR reviews is also detailed.
Related event: Local Qwen3.8-27B Setup Guides(2 posts)→
More from coding & agent
- Obsidian's Smart Chat saves AI thread links and status back into your notes — AINewsletter · 2026-08-24
- Same model, different coding agent: harness choice swings scores from 10/10 to 0/10 — oliver-zehentleitner · 2026-08-24
- Practitioner advice: don't use LLM time savings to produce more mediocre code — tokenbender · 2026-08-24
- Dan Luu: there's no reason for software to be slow anymore — LLMs democratize perf work — tokenbender · 2026-08-24
- Google launches Developer Knowledge API/MCP spanning 19 doc sets, now in gcloud CLI — rseroter · 2026-08-24
- How do you collaboratively manage Markdown context files for AI agents? — ImplementJumpy6494 · 2026-08-24