Every 10-16GB VRAM local LLM ranked: Gemma 4 26B-A4B tops the list, only 200B+ models know your anime
WEREWOLF_BX13 · reddit · 2026-10-12
A Reddit user ranked every LLM runnable on 10-16GB VRAM (26-32GB RAM) for local roleplay/dialogue, scoring speed (≥10 t/s), coherence, multilingual typo handling, slop-phrase frequency, MoE vs full, censorship, and instruction-following (tested with 23 intertwined Pandora rules):
Top picks:
- Gemma 4 26B-A4B — better dialogue/monologue and humor than Qwen 35B-A3B;
- Qwen 3.6 35B-A3B (Abliterated) — fast and accurate but prone to format overfitting;
- Gemma 4 Instruct 19B-A1B Heretic; 4. The Omega Directive 14B (low slop, nice prose); 5. GLM-4.7-Flash 31B-A3B (Abliterated) — 14 models total.
Takeaways: just pick any quant above Q4K; CoT models are a waste of tokens here; and a knowledge quiz found only models above 200B actually carry anime/manga/movie trivia—120B rarely does, and almost nobody can run those.
More from Models
- $/M tokens is broken: Opus 5.5 costs ~6x Sonnet per task at similar scores — lordmairtis · 2026-10-12
- Mathematicians push back on OpenAI's model-generated proofs — ZeeshanZiaML · 2026-10-12
- Anthropic's Claude can't even do substring search in chat history, user shows — AaronBergman18 · 2026-10-12
- Every Recent Claude Loves to Say Things Are 'Carried' or 'Held' — repligate · 2026-10-12
- Claude 3 Opus rolls an "Impossible: Absolute Success" in Disco Elysium-style RP — repligate · 2026-10-12
- Microsoft quietly lists Decision-1, a model that returns calibrated probability scores instead of text — usamawahabkhan · 2026-10-12