20 open-source OCR & PDF extraction tools for RAG pipelines, sorted into 5 categories

MaryamMiradi · x · 2026-10-07

A curated map of 20 open-source OCR + PDF extraction tools for building RAG systems and AI agents over real documents—tables, formulas, scanned pages, multi-column layouts, charts, headers, footnotes—sorted into five categories:

Key takeaway: the best OCR model is not necessarily the best document pipeline. For production agents the goal is preserving enough structure and meaning for the model to reason correctly. Typical pipeline: PDF → native text/layout → OCR/VLM → tables+formulas → reading order → Markdown/JSON → verification → RAG/agent context. A Docling vs MinerU vs Marker vs olmOCR vs PaddleOCR vs OCRFlux benchmark may follow.

Original post →

More from coding & agent

coding & agent channel →