Selecting local coding models for high RAM, limited VRAM setups
Frail_Waif · reddit · 2026-08-30
A user seeks advice on running local agentic coding models on a workstation with 256GB RAM but limited GPU power (RTX A4500 20GB + 940MX 12GB). The debate is between using a quantized Qwen 32B model that fits the VRAM versus a larger model utilizing the massive RAM. The user plans to use WSL2 and llama.cpp, asking if the hardware is sufficient for productive 8-hour workdays to convince colleagues.
More from coding & agent
- Custom Agent panel supports adding archetypes and mid-session intervention — BLUECOW009 · 2026-08-30
- Karpathy's 2-Hour Lecture & Guide to Agent Loops and Graphs — JohnAlexander · 2026-08-30
- Uber's Software Factory: 70% of PRs Attributed to AI Agents — JohnAlexander · 2026-08-30
- Debate on AI Agent Communication: Is Silence More Efficient? — julianharris · 2026-08-30
- GLM-5.3 Released for Agentic Coding at 30% Lower Cost — markjeffrey · 2026-08-30
- Omarchy Multi-Agent Setup: Native MCP Mail and Board — BLUECOW009 · 2026-08-30