OpenDataLoader PDF: open-source parser tops benchmarks with 0.907 accuracy
mdancho84 · x · 2026-09-15
OpenDataLoader PDF is a new open-source parser for AI-ready data that has racked up 29k GitHub stars. It extracts Markdown, HTML, and JSON with bounding boxes from any PDF, handling complex layouts, tables, and nested structures. It runs fully local on CPU, with an AI-hybrid mode for tricky pages. On a benchmark of 200 real-world documents it scores 0.907 overall and 0.928 table accuracy, ranking #1. SDKs available for Python, Java, and Node, free and open source.
Related event: Open-Source PDF Parser OpenDataLoader Tops Benchmarks(2 posts)→
More from coding & agent
- SWE: pre-nerf Opus 4.6 was the goat, and smarter models won't make SWE easier — LouMM · 2026-09-15
- Elyx debuts with a new AI-friendly design file format for designers and agents — michalmalewicz · 2026-09-15
- Hour-long deep dive with an AI agents expert on workflows, skills, and making money — Rasmic · 2026-09-15
- Skill+CLI vs MCP: a developer asks which integration route wins on tokens and reliability — PrinceSauromates · 2026-09-15
- Open-source skill turns one sentence into a single-file web 3D documentary via Claude Code — xiaohu · 2026-09-15
- Hyper3D launches MCP service, letting Codex auto-build 3D models into explainer sites — xiaohu · 2026-09-15