H Company's NeoMME: 260M encoder matches 3.75B ColQwen2.5 on ViDoRe v3 with 14x fewer params

tomaarsen · x · 2026-09-07

H Company open-sourced NeoMME, a family of 260M/800M multilingual multimodal encders built as a single bidirectional Transformer trained from scratch with a masked discrete-diffusion objective — no pretrained vision tower, no causal decoder, 16,384-token context.

Highlights

All checkpoints are Apache 2.0, with a technical report and Visual RAG demo.

Related event: H Company Open-Sources NeoMME: 260M Multimodal Encoder Matches ColQwen2.5(2 posts)→

Original post →

More from Models

Models channel →