New Model's KV Cache QAT and Dropped MTP Spark Debate

Analysts digging into a new model's technical report found its KV cache is quantization-aware trained (explaining strong fp4 performance), MTP was dropped as not worth it, and its engram uses prime-sized tables with 4-gram and fp8 lookups.

2026-09-11 ~ 2026-09-11 · 7 related posts