Reddit thread maps the best tools for running coding agents on 8–12GB local models

salgado18 · reddit · 2026-07-22

The author asks what tools and harnesses people use to run complex coding tasks on small local models that fit in 8–12 GB of VRAM.

They argue that, on constrained models, the harness matters as much as the model itself: persistent memory, code search, call graphs, AST-based navigation, and fewer tool calls can reduce token waste and improve effectiveness.

Tools mentioned include:

Original post →

More from coding & agent

coding & agent channel →