LightOn Open-Sources LightOnOCR-3, Topping OCR Benchmarks in Three Sizes
On October 9, French company LightOn released and open-sourced the third generation of its document intelligence model family, LightOnOCR-3, available in 0.8B, 1B, and 4B parameter sizes. All are open-sourced on Hugging Face under the Apache 2.0 license, with weights, code, and reproducible resources included, making them suitable for both research and commercial use. The models go beyond traditional OCR: a single call can handle text and handwriting recognition, document layout analysis (with labels and bounding boxes), image captioning, and chart data extraction, while preserving document structure. Blogger Igor Carron described its performance as "insane" and opened an online demo; several users suggested testing it against the hardest document samples.
Confirmed
- The release comes from the LightOn team (in collaboration with CavaillesAdrien and others). The release thread publicly details how the grounding data was constructed, as well as the use of intermediate checkpoints from the training process; training data processing details include using LightOnOCR 2 reference text to guide logical chunking and combining multiple layout models for complementary coverage.
- The 0.8B and 4B versions adopt the Qwen3.5 vision-language architecture, supporting visual understanding of image inputs (the 0.8B also supports video), and support empty-prompt operation.
- Benchmarks: on ParseBench, the 4B (75.1) and 0.8B (74.6) take the top two spots; on olmOCR-Bench, the 4B scores 86.3, behind only the 35B Infinity Parser, while leading the Chandra series and outperforming Mistral OCR.
- Previous models have surpassed 4.5 million downloads on Hugging Face.
- Igor Carron demonstrated the model seamlessly extracting hard-to-read German passages from documents in hands-on tests.
Why it matters
- "Good retrieval starts with good parsing (GIGO)" — document parsing quality directly determines the performance of downstream applications such as RAG. LightOnOCR-3 combines OCR, layout analysis, and chart extraction into one small open-source model, lowering the barrier to local deployment and commercial use, and with its small sizes it approaches or even surpasses larger closed-source or high-parameter solutions on the leaderboards.
2026-10-09 ~ 2026-10-09 · 22 related posts
Primary sources
- [source] LightOn Releases LightOnOCR-3: Open-Weight OCR Models Beat Mistral OCR on Benchmarks — wjb_mattingly · 2026-10-09
- [source] LightOnOCR-3 released: 0.8B/1B/4B open models for OCR, layout, charts under Apache 2.0 — antoine_chaffin · 2026-10-09
- LightOn launches LightOnOCR-3, topping open-weight OCR benchmarks over Mistral OCR 4.1 — IgorCarron · 2026-10-09
- LightOnOCR-3 adds visual grounding and image description to document parsing — IgorCarron · 2026-10-09
- France's LightOn ships open-source LightOnOCR-3: 0.8B/4B models top OlmOCR-Bench and ParseBench — IgorCarron · 2026-10-09
- LightOnOCR-3's 0.8B scores 85.5 on olmOCR-Bench, nearly matching a 35B parser — IgorCarron · 2026-10-09
- LightOnOCR-3 takes top two open-weight spots on ParseBench: 4B scores 75.1 — IgorCarron · 2026-10-09
- LightOn Open-Sources LightOnOCR-3 0.8B Vision Document Model Under Apache 2.0 — IgorCarron · 2026-10-09
- LightOnOCR-3 draws praise: try it on your hardest OCR examples — IgorCarron · 2026-10-09
- LightOn releases LightOnOCR-3, shares how intermediate checkpoints improved training data — IgorCarron · 2026-10-09
- Try LightOnOCR-3 on your hardest document: early user demos circulate — IgorCarron · 2026-10-09
- LightOn launches LightOnOCR-3: 0.8B/4B models do OCR, captions and chart extraction in one pass — IgorCarron · 2026-10-09
- LightOnOCR-3 tops ParseBench with 4B and 0.8B models, near-best on olmOCR-Bench — IgorCarron · 2026-10-09
- LightOnOCR-3 released: open-source OCR family in 0.8B/1B/4B with layout grounding and chart extraction — IgorCarron · 2026-10-09
- LightOnOCR-3 OCR Model Available to Try Online, Demo Shared — IgorCarron · 2026-10-09
- [source] LightOnOCR-3 Seamlessly Extracts Hard-to-Read German Text in Demo — IgorCarron · 2026-10-09
- LightOnOCR-3 draws buzz: "insane" says early watcher — IgorCarron · 2026-10-09
5 near-duplicate retellings: IgorCarron · IgorCarron · IgorCarron · IgorCarron · IgorCarron