98,877 historical newspaper pages with OCR text, word boxes released on HF Hub
vanstriendaniel · x · 2026-10-01
A new dataset of 98,877 historical newspaper pages (1700s–1940s), sourced via Europeana, is now on the Hugging Face Hub.
- Each page ships with its original OCR text
- Includes word-level bounding boxes and confidence scores
- Freely available for historical text recognition, OCR evaluation, and document understanding research
Related event: Hugging Face Hosts 98,877-Page Historic Newspaper Dataset(2 posts)→
More from Research
- ctjlewis offers $15,000 in Astra credits to mathematicians — ctjlewis · 2026-10-01
- Gemini 4 Argon beats quantum computing baseline by 40% in minutes, says Google researcher — LinusEkenstam · 2026-10-01
- Qwen3.8-27B-pi: Effort-Ordered Reasoning for Agentic Coding on Hugging Face — paf1138 · 2026-10-01
- State of Clinical AI 2026 report published in BMJ Digital Health & AI — davidjhwu · 2026-10-01
- UChicago brain-machine interfaces let paralyzed patients feel touch via thought — plopesresearch · 2026-10-01
- CLM author responds to RLM debate: context as a variable, and the two can combine into stronger agents — RulinShao · 2026-10-01