Every 10-16GB VRAM local LLM ranked: Gemma 4 26B-A4B tops the list, only 200B+ models know your anime

WEREWOLF_BX13 · reddit · 2026-10-12

A Reddit user ranked every LLM runnable on 10-16GB VRAM (26-32GB RAM) for local roleplay/dialogue, scoring speed (≥10 t/s), coherence, multilingual typo handling, slop-phrase frequency, MoE vs full, censorship, and instruction-following (tested with 23 intertwined Pandora rules):

Top picks:

Takeaways: just pick any quant above Q4K; CoT models are a waste of tokens here; and a knowledge quiz found only models above 200B actually carry anime/manga/movie trivia—120B rarely does, and almost nobody can run those.

Original post →

More from Models

Models channel →