FineBooks Benchmark by HF & EleutherAI: Open OCR Models Hit 97.6% Accuracy on Historical Texts
vanstriendaniel · x · 2026-08-10
Hugging Face and EleutherAI jointly launched the FineBooks project to evaluate the capability of modern open-source OCR models in processing public domain historical literature.
- Background: While millions of historical books have been digitized, early OCR quality is often poor. High-quality public domain text is a crucial data source for training open AI models.
- Benchmark: The team evaluated 14 open-source models based on 2,165 pages of expert human transcription.
- Results: The best model achieved 97.6% character accuracy on historical prints at a highly cost-effective rate (under $2 per 1,000 pages).
This demonstrates that Vision Language Models (VLMs) now have the potential to re-OCR massive historical archives at a low cost, unlocking vast knowledge bases.
More from Research
- Stanford's ChatEHR Deployment: $6M+ Estimated First-Year ROI and Inadequacies of Benchmarks — EricTopol · 2026-08-10
- ICML 2026 Machine Unlearning Tutorial Released with Videos and Slides — thegautamkamath · 2026-08-10
- SceneGen Generates 3D Scenes from a Single Image in One Feedforward Pass — tom_doerr · 2026-08-10
- Surya Ganguli Shares Top Summer Schools in Computational Neuroscience and AI — SuryaGanguli · 2026-08-10
- IFDS Research Expo Aug 10-12 at UW-Madison: AI foundations, ML, stats, optimization — prof_kamilov · 2026-08-10
- Primus Launches Autonomous ML Research Agent: 30x Faster from Hypothesis to Paper — JayAlammar · 2026-08-10