Apple open-sources LensVLM-9B, a VLM that reads compressed text images and selectively expands relevant pages
jacek2023 · reddit · 2026-09-24
Apple released LensVLM-9B, a 9B vision-language model that scans compressed images of text and uses learned tools to selectively expand only the relevant pages to their uncompressed form, saving context. Paper (arXiv:2605.07019) and code (ml-lensvlm) are public, with a GGUF quantization already on Hugging Face. Weights, including Apple's Qwen modifications, ship under the Apple ML Research Model License; source code uses the Apple Sample Code License.
Related event: Apple Open-Sources LensVLM-9B to Save Tokens via Compressed Document Images(2 posts)→
More from Research
- Frontier labs' health week: Opus 5.5 cut 40%, 950 agents find new enzyme — HealthcareAIGuy · 2026-09-24
- Anthropic says Claude uncovered unknown CRISPR-like enzyme system in phage DNA — ccerrato147 · 2026-09-24
- One picture each: gradient descent and backprop explained without arithmetic — alfcnz · 2026-09-24
- Neural net series: multi-layer nets and why cross-entropy and squared error share the same gradient — alfcnz · 2026-09-24
- Teaching Thread: From Perceptrons to Hidden-Layer Representations in Driving — alfcnz · 2026-09-24
- Simulated fly on LSD: connectome wired with eyes shows T4/T5 activity up 10–20% — Merzmensch · 2026-09-24