Peking University Introduces ReBA for Vision-Language MoE Load Balancing

PekingUniversity · hf · 2026-08-04

In Vision-Language Mixture-of-Experts (MoE) models, unbalanced image and text token counts cause standard auxiliary losses to mask true modality-specific load errors. Researchers from Peking University found that a trained router's load imbalance can change more than fivefold across different image resolutions.

To tackle this, the team proposes ReBA (Relax Within, Balance Across). Based on the geometric structure of router inputs—where image and text occupy distinct regions and visual tokens group by source image—ReBA implements two targeted strategies:

Experiments across four split backbones show that ReBA maintains task accuracy comparable to standard methods while significantly reducing load imbalance across various inputs and improving robustness against resolution and tiling shifts. Code is open-sourced.

Original post →

More from Research

Research channel →