Best open models by VRAM: 4B nears 9B-class on 8GB, Qwen 27B tops 24GB
victormustar · x · 2026-09-09
A community guide to picking open models by local VRAM:
8GB VRAM
- Spark-X2.5-4B: benchmarks suggest near-9B performance, focused on coding & agentic use, 1M token context;
- Bonsai-27B: Qwen3.6-27B quantized to Q2 — closest to frontier on 8GB, with 262K context and vision.
16GB VRAM
- Gemma-4-12B: top vision pick, rich world knowledge, great at finding things in images/videos and coherent over long contexts; not the best agent.
24GB VRAM (top pick)
- Qwen3.8-27B, especially Mia Lab's exl3-3.5bpw quantization.
A practical cheat sheet for running open models on your own hardware.
More from Models
- Dev claims turning off thinking mode makes AI models 5x faster, urges default non-thinking modes — Daniel_Farinax · 2026-09-09
- Yoav Goldberg: OpenAI's "cannot rule out" user data use is a lawyered admission — MannyKayy · 2026-09-09
- Heavy user on Astra: beats a junior hire on cost, but still fumbles simple tasks — RachelVT42 · 2026-09-09
- As Models Master Structured Tasks, Creative Writing Keeps Getting Worse — teodorio · 2026-09-09
- Math benchmark success shows log-linear diminishing returns with test-time compute, says ramez — sebkrier · 2026-09-09
- Toby Ord: image models are like a lossy compression format for photos — tobyordoxford · 2026-09-09