Peking University Introduces ReBA for Vision-Language MoE Load Balancing
PekingUniversity · hf · 2026-08-04
In Vision-Language Mixture-of-Experts (MoE) models, unbalanced image and text token counts cause standard auxiliary losses to mask true modality-specific load errors. Researchers from Peking University found that a trained router's load imbalance can change more than fivefold across different image resolutions.
To tackle this, the team proposes ReBA (Relax Within, Balance Across). Based on the geometric structure of router inputs—where image and text occupy distinct regions and visual tokens group by source image—ReBA implements two targeted strategies:
- Applies separate balancing terms for image and text.
- Assigns an equal-weight routing instance per image.
Experiments across four split backbones show that ReBA maintains task accuracy comparable to standard methods while significantly reducing load imbalance across various inputs and improving robustness against resolution and tiling shifts. Code is open-sourced.
More from Research
- Kuaishou's KDD 2026 Paper: Introducing PlatformBid, First Platform-Perspective Auto-Bidding Benchmark — 机器之心 · 2026-08-04
- LeRobot now supports 30+ robot hardware integrations with drop-in plugins — m_wulfmeier · 2026-08-04
- Roasting Academia: NeurIPS is Essentially Just Scrolling OpenReview — abursuc · 2026-08-04
- Open-Source LoRA for Satellite Image Editing Uses AI to Auto-Generate Training Data — LimitlessSaint · 2026-08-04
- Weekend Project: RL-Trained 4B LLM Rewrites AI Text to Fool Open-Source Detectors — matthen2 · 2026-08-04
- NUS Creates Octopus-Inspired Swimming Robot Driven by Just Two Motors — lukas_m_ziegler · 2026-08-04