Reddit asks which local model works best for coding, planning and VS Code workflows
naunen · reddit · 2026-07-27
A Reddit user asks for recommendations on the best local model for coding and thinking after getting tired of paying $200/month for Claude Opus 4.7.
They say they can run a 4-bit GLM 5.2 at about 4 tokens per second and ask whether they should switch to something like Qwen3 Coder 480B or another model. The main use case is building apps, bots and websites in VS Code, with lots of back-and-forth reasoning and planning rather than just one-shot code generation.
More from coding & agent
- Rewriting Vercel CLI in TypeScript and Statically Compiling to 1.28MB — cramforce · 2026-07-27
- “Nobody writes .tikz code anymore,” jokes a developer in the AI era — drscotthawley · 2026-07-27
- ChatGPT app adds thread orchestration tools and sub-agent workflows — pvncher · 2026-07-27
- Buzz agent team goes live with an 11-agent orchestration roster — builditwithjoe · 2026-07-27
- Swarms adds Frenzy Hub, redesigns Marketplace, and ships Gemini 3.6 Flash — KyeGomezB · 2026-07-27
- AI-Trader adds an MCP server so LLMs can run trading backtests — tom_doerr · 2026-07-27