Model reportedly uses QAT for KV cache, explaining strong fp4 performance while dropping MTP
stochasticchasm · x · 2026-09-11
An observer notes the new model applies quantization-aware training (QAT) to the KV cache as well, which explains its better-than-peers performance at fp4 KV cache. The author also noticed MTP (multi-token prediction) was removed, speculating the performance uplift wasn't ultimately worth it.
More from Infra
- DIY-friendly KiCad footprints for AI MELF resistors, milled at home — debreuil · 2026-09-11
- Inference providers barely break even: $10K revenue yields just $200 profit — metalvendetta · 2026-09-11
- Colocated async RL gains steam as observers speculate k3 uses it too — stochasticchasm · 2026-09-11
- Peter Diamandis: The AI race is becoming the biggest construction project of our generation — PeterDiamandis · 2026-09-11
- k3 Report Section Confirms Millions of Concurrent Sandboxes in Its RL Training Run — stochasticchasm · 2026-09-11
- Pentagon in talks to lend roughly $5 billion to AI cloud startup Fluidstack — vitaliychiley · 2026-09-11