Best local coding model for 6GB VRAM: Balancing Qwen performance and response times

BrianScottGregory · reddit · 2026-08-24

A developer seeks recommendations for a reliable local coding model constrained by 6GB VRAM and 64GB RAM. While Qwen 3.8-27B was considered, response times were impractically long (up to an hour). The user, working in C/C++/Python, prefers uncensored models for security research but is open to censored ones if they perform better. The discussion highlights the trade-offs between model size, speed, and hardware limitations.

Original post →

More from coding & agent

coding & agent channel →