Entire arXiv released as Hugging Face dataset: 3.1M papers, every version, 16TB
burny_tech · x · 2026-09-20
secemp9 released the entire arXiv site as a Hugging Face dataset: 3,148,796 papers, every version, in LaTeX, PDFs, PostScript, and HTML — 16TB in total. Called 'an amazing pretraining corpus,' it's a major resource for training and fine-tuning research.
Related event: Full arXiv mirror lands on Hugging Face: 3.14M papers, 16TB(8 posts)→
More from Research
- LangChain's Jev evaluator cuts agent eval score variance by up to 913x at 1/80th the cost — LangChain · 2026-09-20
- Open-source NL logic interpreter unifies facts with Jev, queries cost a quarter cent — narphorium · 2026-09-20
- Annotating CUA agent data: tasks go stale, so pseudo-annotate the tasks themselves — mervenoyann · 2026-09-20
- Science study: kids' brain patterns reflect socioeconomic status, not innate IQ — burny_tech · 2026-09-20
- dQwen3.5: hybrid-attention diffusion LMs hit same loss with half the tokens — burny_tech · 2026-09-20
- New paper: recursive looping boosts pre-training scaling exponents — burny_tech · 2026-09-20