Entire arXiv Published on Hugging Face: 3.14M Papers, 16TB
On September 20, developer secemp9 packaged the entire arXiv site into a public Hugging Face dataset, secemp9/arxiv-complete. This is a complete site-level mirror of arXiv, worth attention from researchers and developers who need large-scale paper corpora.
Confirmed
- The dataset contains 3,148,796 papers, covering every version (including historical revisions), totaling roughly 16TB
- Available in multiple formats: LaTeX source files, PDF, PostScript, and HTML
- Hosting costs about $249 per month, or roughly $3,000 per year; pulling the data from AWS earlier also incurred about $626 in egress fees
- secemp9 found the official arXiv S3 bucket incomplete and had to fill gaps using other sources such as GCS
Why it matters
- This is a rare full-corpus, all-versions, multi-format public arXiv dataset, significantly lowering the barrier for researchers to obtain and reproduce paper corpora
- An independent developer is paying bandwidth and storage costs out of pocket, so sustainability is questionable; he has opened a donation page seeking community help, and long-term maintenance depends on community support
2026-09-20 ~ 2026-09-20 · 7 related posts
Primary sources
- [source] Entire arXiv uploaded to Hugging Face: 3.15M papers, every version, 16TB in LaTeX/PDF/HTML — secemp9 · 2026-09-20
- Entire arXiv Site Released as 16TB Dataset: 3.1M Papers on HuggingFace — secemp9 · 2026-09-20
- [source] Dev pays $249/month to host arXiv dataset mirror, opens donations to cover costs — secemp9 · 2026-09-20
4 near-duplicate retellings: secemp9 · MaziyarPanahi · MaziyarPanahi · burny_tech