NeoMME: single-tower multimodal multilingual encoders beat sub-800M rivals on ViDoRe v3

CShorten30 · x · 2026-09-03

NeoMME is a family of 260M/800M multimodal-native multilingual bidirectional encoders that process text tokens and raw image patches in a single bidirectional Transformer — no pretrained vision tower, text encoder, or decoder — pretrained from scratch with a masked discrete-diffusion objective.

Related event: LightOn Unveils NeoMME: Single-Tower Multimodal Encoder That Leads Sub-800M Retrieval(4 posts)→

Original post →

More from Multimodal

Multimodal channel →