SMAT: merge-aware training lifts merged model scores up to 2.16 with <2% overhead

PolyUHK · hf · 2026-09-29

SMAT frames common model merging methods as three operations from an expert's perspective — Scale (reweighting its own update), Mask (removing selected coordinates), and Perturb (adding other experts' updates) — and jointly optimizes expert loss plus expected loss at simulated merged parameters via sampled scaling coefficients, masks, and noise.

Engineering tricks (periodic scheduling, kernel fusion, parameter storage switching) keep it to one forward and one backward pass per step. Across four language and vision-language backbones, SMAT improves mean score by 1.07–2.16 points over the strongest per-backbone baseline across five merging methods, with under 2% training-time overhead.

Original post →

More from Research

Research channel →