Arch breakdown: dropping convs for Muon optimizer, 3x3 pixel unshuffle for vision
stochasticchasm · x · 2026-09-11
Researchers dissect a newly released model's architecture: multimodality is handled minimally by feeding visual tokens to the backbone, with an aggressive 3x3 pixel unshuffle (vs. the usual 2x2).
Another notable choice is removing conv layers to work with the Muon optimizer—speculation that despite Muon supporting convs via flattening, the team may have found it underperformed in practice.
More from Models
- Microsoft Patches Record 974 Vulnerabilities, Mostly Found by AI — Distinct-Question-16 · 2026-09-11
- A Four-Step Verification Method to Catch AI That Fakes Reading Financial Reports — anthara_ai · 2026-09-11
- DeepSeek V4.1 Flash tops Vals open-weight index at $0.30 per test, with the smallest skills gap — teortaxesTex · 2026-09-11
- Do You Really Need Flagship Models? Dev Argues Medium Effort Covers 80% of Coding — iamaliveix · 2026-09-11
- OpenAI appears to be quietly rolling out managed Agents on its platform — testingcatalog · 2026-09-11
- 30B Open Model OpenResearcher Beats GPT-4.1 on BrowseComp-Plus — TheZachMueller · 2026-09-11