Entire arXiv released as a 16TB HuggingFace dataset: 3.1M papers, every version

udmrzn · x · 2026-09-21

The entire arXiv site has been packaged as a dataset on HuggingFace: 3,148,796 papers, every version, in LaTeX, PDF, PostScript and HTML formats — 16TB in total.

The reposter notes ICLR 2027 submissions already exceed all previous years combined, suggesting this kind of full-corpus release is just the beginning for AI-assisted research at scale.

Related event: Developer publishes entire arXiv as a 16TB dataset on Hugging Face(9 posts)→

Original post →

More from Research

Research channel →