Selecting local coding models for high RAM, limited VRAM setups

Frail_Waif · reddit · 2026-08-30

A user seeks advice on running local agentic coding models on a workstation with 256GB RAM but limited GPU power (RTX A4500 20GB + 940MX 12GB). The debate is between using a quantized Qwen 32B model that fits the VRAM versus a larger model utilizing the massive RAM. The user plans to use WSL2 and llama.cpp, asking if the hardware is sufficient for productive 8-hour workdays to convince colleagues.

Original post →

More from coding & agent

coding & agent channel →