98,877 historical newspaper pages (1700s-1940s) with OCR text and word boxes hit Hugging Face
vanstriendaniel · x · 2026-10-01
Daniel van Strien uploaded a dataset of 98,877 historical newspaper pages (1700s-1940s) to the Hugging Face Hub as biglam/europeananewspapersimages (sourced from Europeana). Each page includes:
- Original OCR text
- Word-level bounding boxes (ALTO format)
- OCR confidence scores
Available in parquet under a public-domain license, spanning German, Estonian, Serbian and 7+ more languages—suited for document layout analysis, image-to-text, and historical document research.
Related event: Hugging Face Hosts 98,877-Page Historic Newspaper Dataset(2 posts)→
More from Research
- GDP.xlsx benchmark: 70 real spreadsheet tasks, best frontier agent scores only 38.3% — echen · 2026-10-01
- Stanford HAI free seminar: Arena CEO Anastasios Angelopoulos on measuring AI in the real world — StanfordHAI · 2026-10-01
- ctjlewis offers $15,000 in Astra credits to mathematicians — ctjlewis · 2026-10-01
- Gemini 4 Argon beats quantum computing baseline by 40% in minutes, says Google researcher — LinusEkenstam · 2026-10-01
- Qwen3.8-27B-pi: Effort-Ordered Reasoning for Agentic Coding on Hugging Face — paf1138 · 2026-10-01
- State of Clinical AI 2026 report published in BMJ Digital Health & AI — davidjhwu · 2026-10-01