Meta's MoEMB: MoE scaling beats 4x-larger embedders with just 3B active params
_reachsumit · x · 2026-09-09
Meta researchers propose MoEMB, scaling universal multimodal embeddings along the expert axis with mixture-of-experts instead of larger dimensions or reasoning tokens. MoE preserves single-vector, non-autoregressive encoding while sidestepping the contrastive-learning batch-size tradeoff. With only 3B active parameters, MoEMB sets new SOTA on MMEB-V2 and MRMR among models trained on public MMEB-family data, surpassing reasoning-token (TTE) methods with >4x active parameters at far lower compute, plus the first comprehensive study of adaptive computation for MoE embedders.
More from Multimodal
- Full AI video pipeline: Qwen 3.8 + Flux 2 Klein + Minimax H3 + LTX 2.5 + Breeze TTS — CQDSN · 2026-09-09
- Minimax H3 videos look h264-compressed regardless of settings, Redditor reports — FoxTrotte · 2026-09-09
- Mage brings MiniMax H3 & H3 Turbo with unlimited video generation, LoRA support — MiniMax_AI · 2026-09-09
- One-line prompt test shows GPT Image 2.5 producing near-indistinguishable iPhone-style photos — Scobleizer · 2026-09-09
- New Flux 2 Klein 9B LoRA gives precise eye-direction control via a red dot — Euphoric_Attorney271 · 2026-09-09
- Four made-up 'imaginary tokens' for Midjourney v8.2 that evoke ghost civilizations and recall fog — LudovicCreator · 2026-09-09