Entire arXiv site released as 16TB Hugging Face dataset: 3.1M papers, all versions

MaziyarPanahi · x · 2026-09-20

A new dataset on Hugging Face covers the entire arXiv site: 3,148,796 papers with every historical version included, shipped in LaTeX, PDF, PostScript, and HTML formats — about 16 TB in total. Commenters joked about the scale, noting you'd practically need a cluster (or GPU-accelerated tooling) to compile that much LaTeX, and thanked the maintainers for the effort.

Related event: Full arXiv mirror lands on Hugging Face: 3.14M papers, 16TB(8 posts)→

Original post →

More from Research

Research channel →