Which local LLM is best for a 36GB VRAM coding agent setup?
hovikyan · reddit · 2026-07-28
The poster has about 36GB of VRAM and wants a local LLM to drive an agent like Hermes or OpenCode on software projects and other personal tasks. They ask which of several locally runnable candidates—Qwen3.6 27B, Qwen3.6 35B A3B, Gemma 4 31B, GLM 4.7 Flash, or Llama 3.3 70B—would be best for that setup and use case.
Related event: Selecting Local Agent Models for 36GB VRAM(2 posts)→
More from coding & agent
- AI won’t shrink software demand—it may unlock a much larger services market — joecole · 2026-07-28
- A practical cloud-agent stack: Devin, Cursor, Open-Inspect, Infisical and more — vinvan · 2026-07-28
- Stripe upgrades Directory with instant listings and MCP search for agents — jeff_weinstein · 2026-07-28
- AIWayfinder Launches DCA v0.6.0 with Zero-Intervention Self-Healing Daemon — templecrash · 2026-07-28
- Council 1.2 adds blind multi-model review and a guest seat for any external answer — ahumanbeingmars · 2026-07-28
- A company with 2 H200s asks which coding model and vLLM setup can serve 4–10 users — redblood252 · 2026-07-28