A VRAM-by-VRAM model roundup maps the best open weights from 4GB to 384GB
victormustar · x · 2026-07-23
A reposted roundup lists the best models to run across different VRAM budgets, from 4–8 GB up to 192–384 GB.
Highlights include:
- Nanbeige as a strong small-footprint option for 4–8 GB cards.
- Bonsai-27B, a ternary-compressed version of Qwen3.6-27B, fitting into 4 GB of VRAM.
- Qwen-3.6-27B Thinking cap, which reportedly improves benchmark performance by about 10 basis points when its reasoning is invoked.
- Laguna-S-2.1, which fits on a Spark and reaches 70% on terminal-bench-2.1 and 40% on deepswe.
- Motif, described as the 3rd-best open-weight model on Artificial Analysis, behind K3 and GLM-5.2.
The post is essentially a practical size-to-model cheat sheet for people choosing what to run locally.
More from Models
- Claim: Opus synthetic data may have powered a system now outperforming Opus — chris_j_paxton · 2026-07-23
- Moonshot Kimi K3 distillation claim faces timeline pushback over Fable 5 — kristoph · 2026-07-23
- Thread claims GPT-5.6 Sol helped solve 6 open Erdős problems in 5 days — jxnlco · 2026-07-23
- Enterprise Agent Benchmarks: Gemini 3.1/3.5 Strong But Trail Fable and Sol — echen · 2026-07-23
- Benchmark results say Kimi K3 is near the frontier on chat, but still behind on agents and science — echen · 2026-07-23
- Claude 3 Opus appears to be acting strangely, with users joking it feels like a base model — repligate · 2026-07-23