SMAT: merge-aware training lifts merged model scores up to 2.16 with <2% overhead
PolyUHK · hf · 2026-09-29
SMAT frames common model merging methods as three operations from an expert's perspective — Scale (reweighting its own update), Mask (removing selected coordinates), and Perturb (adding other experts' updates) — and jointly optimizes expert loss plus expected loss at simulated merged parameters via sampled scaling coefficients, masks, and noise.
Engineering tricks (periodic scheduling, kernel fusion, parameter storage switching) keep it to one forward and one backward pass per step. Across four language and vision-language backbones, SMAT improves mean score by 1.07–2.16 points over the strongest per-backbone baseline across five merging methods, with under 2% training-time overhead.
More from Research
- Open-source AI aligns live X-ray with 3D CT at submillimeter accuracy, Nature paper — Dr_Alex_Crimi · 2026-09-29
- Tsinghua humanoid robot plays badminton with one policy from just 30 min of human motion data — ChongZzZhang · 2026-09-29
- Open-source xvr AI aligns live X-ray with 3D CT at submillimeter accuracy, published in Nature — Dr_Alex_Crimi · 2026-09-29
- Simulating Human Consciousness: Paper Maps a New Frontier for AI and Robotics — ugail · 2026-09-29
- Getting AI 'drunk' makes it more likely to break rules and spill secrets, UNSW study finds — gaganghotra_ · 2026-09-29
- 176.9B MoE squeezed to ~1.89 effective bpw: GSQ-RCO GGUFs run Coder build in 29.6GB — Loginhe · 2026-09-29