Full arXiv corpus now on Hugging Face: 3.1M papers, all versions, 16 TB
MaziyarPanahi · x · 2026-09-20
The entire arXiv site has been released as a dataset on Hugging Face: 3,148,796 papers with every version, provided in LaTeX, PDF, PostScript and HTML formats, totaling 16 TB.
Related event: Full arXiv mirror lands on Hugging Face: 3.14M papers, 16TB(8 posts)→
More from Research
- LangChain's Jev evaluator cuts agent eval score variance by up to 913x at 1/80th the cost — LangChain · 2026-09-20
- Open-source NL logic interpreter unifies facts with Jev, queries cost a quarter cent — narphorium · 2026-09-20
- Annotating CUA agent data: tasks go stale, so pseudo-annotate the tasks themselves — mervenoyann · 2026-09-20
- Science study: kids' brain patterns reflect socioeconomic status, not innate IQ — burny_tech · 2026-09-20
- dQwen3.5: hybrid-attention diffusion LMs hit same loss with half the tokens — burny_tech · 2026-09-20
- New paper: recursive looping boosts pre-training scaling exponents — burny_tech · 2026-09-20