LightOn ships NeoMME: natively multimodal encoders at 260M/800M with no vision tower
antoine_chaffin · x · 2026-09-03
LightOn releases NeoMME, rejecting the common practice of using massive generative VLMs as representation models. Instead, the team trained encoders from scratch that are natively text-and-image, multilingual, and built for speed from the start — at just 260M and 800M parameters, with no vision tower.
- Also fine-tuned NeoMME-Retriever, a visual document retriever built on top
- The project delivers on Aurelien Lemonier's long-standing pitch of building a "multimodal ModernBERT"
More from Multimodal
- Claude writes three.js code, Higgsfield renders it: text-to-3D architecture pipeline — heypearlai · 2026-09-03
- Meta's Alexandr Wang touts 'Muse Spark 1.3' voxel planet demo: it's like the Little Prince — alexandr_wang · 2026-09-03
- Seedance 2.5 turns a supermarket complaint into a AAA game boss fight, full prompt shared — azed_ai · 2026-09-03
- Seedance 2.5 demo turns a retail complaint into a AAA game boss fight — azed_ai · 2026-09-03
- How Squad made its launch video with Revid CLI and 98 script revisions — tibo_maker · 2026-09-03
- Snap a photo, drop it into Blender 3D: the two-step photo-to-3D trick — sidahuj · 2026-09-03