Best Models for Local Code Agents on 20GB VRAM
kirisoraa · reddit · 2026-07-13
The poster is looking for models suitable for a local coding agent running on a 64GB DDR5 RAM laptop + 20GB 7900 XT eGPU.
Their current experience:
- Qwen3.6 35B A3B performs best under CPU offload scenarios
- It can run at Q8, 100k context
- It also works with ud-q6kl, full 262k context
- Although a bit slow, 20+ tokens/s is still usable for them
They want to know:
- What other MoE models are worth trying on similar hardware
- Whether to stick with lower-quant but larger models, or choose smaller dense models fully loaded into VRAM
More from Infra
- NeurIPS 2026 workshop will focus on on-device intelligence and local execution — YiMaTweets · 2026-07-21
- How to build a PostgreSQL-backed semantic search pipeline with pgvector and Ollama — KhuyenTran16 · 2026-07-21
- NeurIPS 2026 workshop calls papers on on-device intelligence — YiMaTweets · 2026-07-21
- Milled from Solid Aluminum: AI Rig Multi-GPU Case for Local Compute — dee_hw · 2026-07-21
- FutureCaribbean’s Buildathon offers $50K, H200 compute, and an NYSE pitch — HeyAmit_ · 2026-07-21
- A new series tests which data-science workflows can run on GPUs today — pandeyparul · 2026-07-21