New Model's KV Cache QAT and Dropped MTP Spark Debate
Analysts digging into a new model's technical report found its KV cache is quantization-aware trained (explaining strong fp4 performance), MTP was dropped as not worth it, and its engram uses prime-sized tables with 4-gram and fp8 lookups.
2026-09-11 ~ 2026-09-11 · 7 related posts
- Engram section analysis: prime-sized tables, 4-grams, and fp8 lookup tables — stochasticchasm · 2026-09-11
- MTP dropped in new model's tech report: fp8 lookup tables and prime table sizes noted — stochasticchasm · 2026-09-11
- Model reportedly uses QAT for KV cache, explaining strong fp4 performance while dropping MTP — stochasticchasm · 2026-09-11
- KV cache gets QAT too: why this model beats others at fp4 KV cache — stochasticchasm · 2026-09-11
- Another Model Shifts to Muon Optimizer as It Emerges as the Training Default — stochasticchasm · 2026-09-11
- Muon optimizer becomes the default as new models drop Adam entirely — stochasticchasm · 2026-09-11
- No Input/Output Optimized Params, and No Weight Decay on the LM Head — stochasticchasm · 2026-09-11