1.47M newspaper pages dataset released to make historical records searchable
leppert · x · 2026-09-02
The Institutional Data Initiative and Boston Public Library released a 1.47M-page historical newspaper dataset plus a processing pipeline that breaks pages into their component parts (columns, headlines, ads), fixing collapsed layouts and lost reading order, boosting OCR accuracy and computational addressability.
More from Research
- Prompt optimization may not need search trees: NPO paper shows teacher quality beats elaborate search — rohanpaul_ai · 2026-09-02
- Schmidhuber revisits award-winning NLSOM paper on economies of minds — SchmidhuberAI · 2026-09-02
- MENO hybrid physics-AI simulation cuts plasma computation from 5 days to 4 hours — bravo_abad · 2026-09-02
- Jasper AI open-sources a full cookbook to train a text-to-image model from scratch — dh7net · 2026-09-02
- 2021's CABiNet beats YOLO26-sem on UAVid: +2.7 mIoU at 3x lower latency — Naive-Explanation940 · 2026-09-02
- Research agenda proposed for 'generative cryptography': AI writing crypto protocols — DavideCrapis · 2026-09-02