20 open-source OCR & PDF extraction tools for RAG pipelines, sorted into 5 categories
MaryamMiradi · x · 2026-10-07
A curated map of 20 open-source OCR + PDF extraction tools for building RAG systems and AI agents over real documents—tables, formulas, scanned pages, multi-column layouts, charts, headers, footnotes—sorted into five categories:
- Native PDF extraction: PyMuPDF4LLM, pdfplumber, pypdf—best when the PDF already has clean digital text
- OCR engines: Tesseract, PaddleOCR, Surya, EasyOCR, docTR
- Document parsers: Docling, MinerU, Marker, Unstructured
- VLM-based parsing: olmOCR, OCRFlux, MonkeyOCR, dots.mocr, Zerox—the category the author finds most interesting, where models interpret document structure, not just characters
- Specialist tools: GROBID (scientific papers), Camelot (tables), Nougat (academic equations)
Key takeaway: the best OCR model is not necessarily the best document pipeline. For production agents the goal is preserving enough structure and meaning for the model to reason correctly. Typical pipeline: PDF → native text/layout → OCR/VLM → tables+formulas → reading order → Markdown/JSON → verification → RAG/agent context. A Docling vs MinerU vs Marker vs olmOCR vs PaddleOCR vs OCRFlux benchmark may follow.
More from coding & agent
- Dev open-sources AutoType, a free Wispr Flow alternative for Linux voice typing — premakin · 2026-10-07
- Running two Muse AI agents as a boss-worker digital production factory — NickPassig · 2026-10-07
- Coinbase's Lincoln Murr on paying AI agents via the x402 protocol — MurrLincoln · 2026-10-07
- Resend logs 3M MCP calls last month, up ~30x since April — dsp_ · 2026-10-07
- Dev vibes-clones his 2-year game with Claude: looks right, plays wrong — IanArawjo · 2026-10-07
- Y Combinator-backed Linc launches discovery layer that turns enterprise tribal knowledge into SOPs — ycombinator · 2026-10-07