Local 27B Agent on 12GB VRAM: Qwen3.8 config and long-context practice

PyaesoneP · reddit · 2026-08-24

The author shares practical experience running Qwen 3.8 27B (UDQ4KXL) for local agent coding on an RTX 5070 Ti Mobile (12GB). The setup includes 100K context, specific llama-server launch parameters, and inference settings. To handle context overflow caused by verbose reasoning, the author adopted Magic Context over native compaction, scaling sessions to 3.7M processed tokens and shipping end-to-end features. A hybrid workflow using Claude Opus for final PR reviews is also detailed.

Related event: Local Qwen3.8-27B Setup Guides(2 posts)→

Original post →

More from coding & agent

coding & agent channel →