Entire arXiv site released as 16TB Hugging Face dataset: 3.1M papers, all versions
MaziyarPanahi · x · 2026-09-20
A new dataset on Hugging Face covers the entire arXiv site: 3,148,796 papers with every historical version included, shipped in LaTeX, PDF, PostScript, and HTML formats — about 16 TB in total. Commenters joked about the scale, noting you'd practically need a cluster (or GPU-accelerated tooling) to compile that much LaTeX, and thanked the maintainers for the effort.
Related event: Full arXiv mirror lands on Hugging Face: 3.14M papers, 16TB(8 posts)→
More from Research
- LangChain's Jev evaluator cuts agent eval score variance by up to 913x at 1/80th the cost — LangChain · 2026-09-20
- Open-source NL logic interpreter unifies facts with Jev, queries cost a quarter cent — narphorium · 2026-09-20
- Annotating CUA agent data: tasks go stale, so pseudo-annotate the tasks themselves — mervenoyann · 2026-09-20
- Science study: kids' brain patterns reflect socioeconomic status, not innate IQ — burny_tech · 2026-09-20
- dQwen3.5: hybrid-attention diffusion LMs hit same loss with half the tokens — burny_tech · 2026-09-20
- New paper: recursive looping boosts pre-training scaling exponents — burny_tech · 2026-09-20