NeoMME: single-tower multimodal encoder leads sub-800M models on ViDoRe v3 visual retrieval

_reachsumit · x · 2026-09-03

NeoMME is a family of 260M/800M bidirectional encoders processing multilingual text and raw image patches in one Transformer, pretrained from scratch with a masked discrete-diffusion objective and 16K-token context. Fine-tuned retrievers hit 0.523 (260M) and 0.556 (800M) nDCG@10 on ViDoRe v3, beating all evaluated models under 800M parameters, with 2x encoding throughput of ColModernVBERT on an L40S. Weights are on Hugging Face.

Original post →

More from Multimodal

Multimodal channel →