Entire arXiv released as Hugging Face dataset: 3.1M papers, every version, 16TB

burny_tech · x · 2026-09-20

secemp9 released the entire arXiv site as a Hugging Face dataset: 3,148,796 papers, every version, in LaTeX, PDFs, PostScript, and HTML — 16TB in total. Called 'an amazing pretraining corpus,' it's a major resource for training and fine-tuning research.

Related event: Full arXiv mirror lands on Hugging Face: 3.14M papers, 16TB(8 posts)→

Original post →

More from Research

Research channel →