LightOn Open-Sources LightOnOCR-3: All-in-One Document AI Model in Three Sizes
French company LightOn released the document intelligence model family LightOnOCR-3 on October 9, going beyond traditional OCR: a single model handles text and handwriting recognition, document layout analysis, image captioning, and chart data extraction. The models come in 0.8B, 1B, and 4B parameter sizes, all open-sourced under Apache 2.0 with commercial use allowed.
Confirmed
- All three sizes (0.8B, 1B, 4B) are open-sourced under Apache 2.0 with commercial use permitted (multiple posts consistently relay the official release).
- A single model covers four task types: OCR transcription, layout analysis, image captioning, and chart extraction.
- Per @wjbmattingly, the 0.8B and 4B versions build on the Qwen3.5 vision-language architecture and support empty-prompt usage.
Unconfirmed
- Benchmark details appear only in posts from @IgorCarron and @wjbmattingly, claiming it leads Chandra-O… (text truncated) on OlmOCR-Bench and beats Mistral OCR; exact scores and comparison targets are incomplete — refer to the official evaluation page.
Why it matters
- Document processing pipelines previously required chaining multiple models for OCR, layout analysis, and chart parsing; LightOnOCR-3 unifies these capabilities in one open-source model, cutting deployment and integration costs. Apache 2.0 plus commercial licensing makes it friendly for enterprises and RAG/document-intelligence applications.
2026-10-09 ~ 2026-10-09 · 5 related posts
Primary sources
- [source] LightOn Releases LightOnOCR-3: Open-Weight OCR Models Beat Mistral OCR on Benchmarks — wjb_mattingly · 2026-10-09
- [source] LightOnOCR-3 released: 0.8B/1B/4B open models for OCR, layout, charts under Apache 2.0 — antoine_chaffin · 2026-10-09
- LightOn launches LightOnOCR-3, topping open-weight OCR benchmarks over Mistral OCR 4.1 — IgorCarron · 2026-10-09
2 near-duplicate retellings: IgorCarron · IgorCarron