Model reportedly uses QAT for KV cache, explaining strong fp4 performance while dropping MTP

stochasticchasm · x · 2026-09-11

An observer notes the new model applies quantization-aware training (QAT) to the KV cache as well, which explains its better-than-peers performance at fp4 KV cache. The author also noticed MTP (multi-token prediction) was removed, speculating the performance uplift wasn't ultimately worth it.

Related event: Muon becomes the default optimizer as new model drops MTP and applies QAT to KV cache(7 posts)→

Original post →

More from Infra

Infra channel →