Entire arXiv released as a HuggingFace dataset: 3.1M papers, 16TB in LaTeX, PDF, HTML

secemp9 · x · 2026-09-20

A developer has packaged the entire arXiv site as a public dataset on Hugging Face: 3,148,796 papers with every version included, in LaTeX, PDF, PostScript, and HTML — 16 TB in total. The author notes dedup was version-based, so counting versions brings the total to over 6M documents. A rare full-corpus, multi-format resource for large-scale training, retrieval, or literature analysis.

Related event: Full arXiv mirror lands on Hugging Face: 3.14M papers, 16TB(8 posts)→

Original post →

More from Research

Research channel →