Apple open-sources LensVLM-9B, a VLM that reads compressed text images and selectively expands relevant pages

jacek2023 · reddit · 2026-09-24

Apple released LensVLM-9B, a 9B vision-language model that scans compressed images of text and uses learned tools to selectively expand only the relevant pages to their uncompressed form, saving context. Paper (arXiv:2605.07019) and code (ml-lensvlm) are public, with a GGUF quantization already on Hugging Face. Weights, including Apple's Qwen modifications, ship under the Apple ML Research Model License; source code uses the Apple Sample Code License.

Related event: Apple Open-Sources LensVLM-9B to Save Tokens via Compressed Document Images(2 posts)→

Original post →

More from Research

Research channel →