Meta's MoEMB: MoE scaling beats 4x-larger embedders with just 3B active params

_reachsumit · x · 2026-09-09

Meta researchers propose MoEMB, scaling universal multimodal embeddings along the expert axis with mixture-of-experts instead of larger dimensions or reasoning tokens. MoE preserves single-vector, non-autoregressive encoding while sidestepping the contrastive-learning batch-size tradeoff. With only 3B active parameters, MoEMB sets new SOTA on MMEB-V2 and MRMR among models trained on public MMEB-family data, surpassing reasoning-token (TTE) methods with >4x active parameters at far lower compute, plus the first comprehensive study of adaptive computation for MoE embedders.

Original post →

More from Multimodal

Multimodal channel →