Muon becomes the default optimizer as new model drops MTP and applies QAT to KV cache

stochasticchasm 研读某新模型的技术报告后,总结出多项值得关注的架构与训练设计:该模型在训练中按注意力头切换使用 Muon 优化器(head-wise Muon),完全没有使用 Adam,作者据此判断 Muon 正在成为训练新模型的默认选择。同时,模型的多 token 预测(MTP)被移除,作者猜测是性能提升最终未能证明其价值。

已确认

为什么重要

2026-09-11 ~ 2026-09-11 · 7 related posts

Primary sources