Reddit asks which local model works best for coding, planning and VS Code workflows

naunen · reddit · 2026-07-27

A Reddit user asks for recommendations on the best local model for coding and thinking after getting tired of paying $200/month for Claude Opus 4.7.

They say they can run a 4-bit GLM 5.2 at about 4 tokens per second and ask whether they should switch to something like Qwen3 Coder 480B or another model. The main use case is building apps, bots and websites in VS Code, with lots of back-and-forth reasoning and planning rather than just one-shot code generation.

Original post →

More from coding & agent

coding & agent channel →