NeoMME: single-tower multimodal multilingual encoders beat sub-800M rivals on ViDoRe v3
CShorten30 · x · 2026-09-03
NeoMME is a family of 260M/800M multimodal-native multilingual bidirectional encoders that process text tokens and raw image patches in a single bidirectional Transformer — no pretrained vision tower, text encoder, or decoder — pretrained from scratch with a masked discrete-diffusion objective.
- 16,384-token context, enough to encode two 4K UHD images
- With joint dense + late-interaction heads, NeoMME-Retriever 260M hits 0.523 nDCG@10 on ViDoRe v3, beating all evaluated models below 800M; the 800M version reaches 0.556
- At matched 2048x2048 input on an NVIDIA L40S, the 260M model encodes pages with 2x the throughput of ColModernVBERT
- Hierarchical token pooling and asymmetric quantization compress late-interaction document embeddings for cheaper retrieval
More from Multimodal
- Seedance 2.5 turns a supermarket complaint into a AAA game boss fight, full prompt shared — azed_ai · 2026-09-03
- Seedance 2.5 demo turns a retail complaint into a AAA game boss fight — azed_ai · 2026-09-03
- How Squad made its launch video with Revid CLI and 98 script revisions — tibo_maker · 2026-09-03
- Snap a photo, drop it into Blender 3D: the two-step photo-to-3D trick — sidahuj · 2026-09-03
- Topaz Brings Video Enhancement to the Browser: Upscaling, Frame Interpolation, SDR-to-HDR Without Any App — umesh_ai · 2026-09-03
- ComfyUI Browser UI Keeps Crashing on Heavy Video Workflows, Author Migrated to Desktop — Suspicious_Pizza9529 · 2026-09-03