NeoMME: From-Scratch Multimodal Encoders at 260M/800M With No Vision Tower
CShorten30 · x · 2026-09-03
Tony Wu and Aurelien Lucet released NeoMME, encoders trained from scratch — natively multimodal, multilingual, and built for speed, at 260M and 800M sizes with no vision tower — plus NeoMME-Retriever, a fine-tuned visual document retriever built on top. An interactive demo is available; LFM2.5-VL-450M serves as the VLM component.
More from Research
- 15.9M-parameter OCR model Kraken PP-OCRv6 beats VLMs hundreds of times larger — vanstriendaniel · 2026-09-03
- Ben Recht's forecasting lecture: the math of turning past frequencies into future odds — beenwrekt · 2026-09-03
- Insilico puts Liquid AI's 2.6B model into MMAI Gym, setting SOTA in retrosynthesis — helloiamleonie · 2026-09-03
- We need a better taxonomy for "continual learning": five mechanisms, five tradeoffs — Typical-Scene-5794 · 2026-09-03
- SupraLabs open-sources 5M-param GatedDeltaNet trained on just 2B tokens — LH-Tech_AI · 2026-09-03
- Nature Medicine Matters Arising disputes claims that general LLMs beat clinical AI tools — anshulkundaje · 2026-09-03