Local Test: Nanbeige4.2-3B Lags Behind Qwen MoE in KV Cache Efficiency

TechTefa · reddit · 2026-07-28

A developer tested the quantized Nanbeige4.2-3B on a 4GB VRAM setup. The results show poor KV Cache efficiency, supporting only 6k context in f16 and 12k in q80.

In contrast, he notes that Qwen 3.6 MoE can run a 64k f16 context in just 1.2GB of VRAM. For low-end hardware, the author concludes that sticking to compact Qwen models remains the better choice.

Original post →

More from Models

Models channel →