Arch breakdown: dropping convs for Muon optimizer, 3x3 pixel unshuffle for vision

stochasticchasm · x · 2026-09-11

Researchers dissect a newly released model's architecture: multimodality is handled minimally by feeding visual tokens to the backbone, with an aggressive 3x3 pixel unshuffle (vs. the usual 2x2).

Another notable choice is removing conv layers to work with the Muon optimizer—speculation that despite Muon supporting convs via flattening, the team may have found it underperformed in practice.

Original post →

More from Models

Models channel →