LlamaIndex keeps ParseBench and ExtractBench fully open, taking a swipe at rivals blocking competitors
llama_index · x · 2026-09-05
LlamaIndex founder Jerry Liu says that unlike other OCR players, the company doesn't block competitors from accessing its platform, and remains committed to open, reproducible benchmarks for document parsing and extraction.
- ParseBench (parsing) and ExtractBench (extraction) are fully open, maintained benchmarks, continuously adding new models like 3.8 flash, fable 5.1, and r-1
- The goal is to show the industry how the accuracy-vs-cost pareto frontier evolves
- The team admits mistakes and incorporates community feedback to improve
- Both benchmarks offer GitHub/HuggingFace access
The post doubles as promotion for LlamaIndex's production OCR offering.
More from Models
- GPT-6 Astra tops Terminal Bench 4.0 at half the cost of #2 — charliermarsh · 2026-09-05
- Users report GPT-6 Astra keeps forgetting it can use computer and Gmail MCP tools — Soft_Hand_1971 · 2026-09-05
- Eric Horvitz: Astra's model card shows CoT-based abuse monitoring is getting harder — erichorvitz · 2026-09-05
- GPT-6 Astra lands day-zero on Databricks, touting SOTA agentic reasoning and document processing — matei_zaharia · 2026-09-05
- Unverified: 'Astra' model explodes Runescape bench scores, records on 10/16 skills — scaling01 · 2026-09-05
- The model hyped as AGI two months ago vs. an average GPT-6 Astra output — aidan_mclau · 2026-09-05