With 4GB of VRAM, which small coding model still works best for agentic coding?

GamerWael · reddit · 2026-07-27

A user with only 4GB of VRAM and 40GB of RAM is asking which small model is currently the best fit for agentic coding.

They mention seeing strong buzz around Gemma 4 and Qwen 3.6, but those models are out of reach for their hardware. The open question is which smaller model makes the most sense, and whether llama.cpp is still the right runtime for this setup.

Original post →

More from coding & agent

coding & agent channel →