Qwen3.8-27B on a 7900XTX hits 40 tok/s with 240K context for local agentic coding

W61k3r · reddit · 2026-09-20

A Reddit user shared a full recipe for running Qwen3.8-27B (Q4KM) locally on a single 7900XTX for agentic coding, hitting 40 tok/s with a 240K context window.

Key details:

The user hands off coding tasks to the agent and even had the model generate a control panel from scratch to manage his inference servers. Highly replicable for local coding-agent setups on consumer GPUs.

Original post →

More from coding & agent

coding & agent channel →