LightOnOCR-3 training: using OCR reference text to guide logical block grouping
IgorCarron · x · 2026-10-09
The LightOn team shares a LightOnOCR-3 training-data detail: the goal was to produce logical blocks, not just layout boxes. They used LightOnOCR 2 output as a consistent reference text, with its paragraph boundaries guiding how detected lines were grouped into boxes — creating coherent blocks for downstream chunking. A concrete look at the engineering behind document-parsing training data.
More from Research
- Solo Dev Pretrains 565M Hybrid LLM From Scratch on a Single RTX 4090 — BLUECOW009 · 2026-10-09
- BABA-is-AI: 2024 ICML benchmark that broke SOTA LLMs deserves a 2026 retest — moschles · 2026-10-09
- NVIDIA open-sources NV-Reason-CT, a native 3D vision-language model for CT scans — NVIDIA Developer · 2026-10-09
- One Epoch of Toloka's Enterprise RL Data Boosts Qwen3.5-27B Agent Benchmarks by up to 44pp — MParakhin · 2026-10-09
- srush builds Jax-Lean transpiler to formally verify JAX tensor code — srush_nlp · 2026-10-09
- COLM2026 talk: LLM factual generation-verification gaps evolve across fact lifecycle — caglarml · 2026-10-09