Entire arXiv released as a 16TB HuggingFace dataset: 3.1M papers, every version
udmrzn · x · 2026-09-21
The entire arXiv site has been packaged as a dataset on HuggingFace: 3,148,796 papers, every version, in LaTeX, PDF, PostScript and HTML formats — 16TB in total.
The reposter notes ICLR 2027 submissions already exceed all previous years combined, suggesting this kind of full-corpus release is just the beginning for AI-assisted research at scale.
Related event: Developer publishes entire arXiv as a 16TB dataset on Hugging Face(9 posts)→
More from Research
- Framingham and TCGA were pivotal without AI; models in the loop could boost data generation — anshulkundaje · 2026-09-21
- Schulman's goal-driven research advice sparks debate on method-driven research traps — mattturck · 2026-09-21
- China releases first-round post-quantum crypto candidates across three algorithm categories — matthew_d_green · 2026-09-21
- Tsinghua's DiffuTester generates unit tests with diffusion LLMs 2-3x faster — jiqizhixin · 2026-09-21
- Jev, built by an RLHF/InstructGPT veteran, targets machine decisions with new RLCD training — MaryamMiradi · 2026-09-21
- Cryptographer Matthew Green questions the point of AI-invented crypto research — matthew_d_green · 2026-09-21