NeoMME: single-tower multimodal encoder leads sub-800M models on ViDoRe v3 visual retrieval
_reachsumit · x · 2026-09-03
NeoMME is a family of 260M/800M bidirectional encoders processing multilingual text and raw image patches in one Transformer, pretrained from scratch with a masked discrete-diffusion objective and 16K-token context. Fine-tuned retrievers hit 0.523 (260M) and 0.556 (800M) nDCG@10 on ViDoRe v3, beating all evaluated models under 800M parameters, with 2x encoding throughput of ColModernVBERT on an L40S. Weights are on Hugging Face.
More from Multimodal
- SimLoss Enables Single-Pass Fine-Grained Image Captioning at Multi-Stage Quality — Suryaansh Jain · 2026-09-03
- Joke About Stacking 'DLSS5' on MiniMax H3 Turbo Real-Time AI Video — AIandDesign · 2026-09-03
- AI-Generated Action Short Recreates Baki vs Luke Fight Cinematics — Ok-Vegetable-2455 · 2026-09-03
- RunPod + Minimax H3 produces gibberish speech in ComfyUI; new user asks why — vscience · 2026-09-03
- UNREEL: open-source AI streaming service generates video live, rendering faster than playback — EAccelerate_42 · 2026-09-03
- ComfyUI workflow generates 3-view character sheets from face+outfit refs — Sea-Advantage-4063 · 2026-09-03