NeoMME: From-Scratch Multimodal Encoders at 260M/800M With No Vision Tower

CShorten30 · x · 2026-09-03

Tony Wu and Aurelien Lucet released NeoMME, encoders trained from scratch — natively multimodal, multilingual, and built for speed, at 260M and 800M sizes with no vision tower — plus NeoMME-Retriever, a fine-tuned visual document retriever built on top. An interactive demo is available; LFM2.5-VL-450M serves as the VLM component.

Related event: LightOn and H Release NeoMME, a Single-Tower Native Multimodal Multilingual Encoder(5 posts)→

Original post →

More from Research

Research channel →