Entire arXiv released as a HuggingFace dataset: 3.1M papers, 16TB in LaTeX, PDF, HTML
secemp9 · x · 2026-09-20
A developer has packaged the entire arXiv site as a public dataset on Hugging Face: 3,148,796 papers with every version included, in LaTeX, PDF, PostScript, and HTML — 16 TB in total. The author notes dedup was version-based, so counting versions brings the total to over 6M documents. A rare full-corpus, multi-format resource for large-scale training, retrieval, or literature analysis.
Related event: Full arXiv mirror lands on Hugging Face: 3.14M papers, 16TB(8 posts)→
More from Research
- LangChain's Jev evaluator cuts agent eval score variance by up to 913x at 1/80th the cost — LangChain · 2026-09-20
- Open-source NL logic interpreter unifies facts with Jev, queries cost a quarter cent — narphorium · 2026-09-20
- Annotating CUA agent data: tasks go stale, so pseudo-annotate the tasks themselves — mervenoyann · 2026-09-20
- Science study: kids' brain patterns reflect socioeconomic status, not innate IQ — burny_tech · 2026-09-20
- dQwen3.5: hybrid-attention diffusion LMs hit same loss with half the tokens — burny_tech · 2026-09-20
- New paper: recursive looping boosts pre-training scaling exponents — burny_tech · 2026-09-20