Liquid AI Launches LFM2.5-VL-3B Lightweight Vision Model, Outperforming Larger Rivals

helloiamleonie · x · 2026-08-12

Liquid AI has officially released LFM2.5-VL-3B, a lightweight vision-language model. Built on the LFM2.5-2.6B base and a SigLIP2 400M NaFlex vision encoder, the model was pre-trained on approximately 34T tokens.

The model excels at reading digital screens across mobile, web, and desktop, grounding objects to coordinates, reading text and charts, and calling tools from text or image inputs. Despite its compact 3B parameter size, it achieves strong benchmark scores: 80.7 on ScreenSpot-v2 and 73.1 on RealWorldQA, outperforming larger models like Gemma-4-E4B and InternVL-3.5-4B. In practical tests, it successfully parsed complex historical manuscripts featuring handwriting and equations.

Related event: LiquidAI Releases LFM2.5-VL-3B Edge Multimodal Model(4 posts)→

Original post →

More from Models

Models channel →