Entire arXiv uploaded to Hugging Face: 3.15M papers, every version, 16TB in LaTeX/PDF/HTML
secemp9 · x · 2026-09-20
User secemp9 released secemp9/arxiv-complete on Hugging Face: the entire arXiv site as a dataset with 3,148,796 papers covering every version, 16 TB total.
- All formats: LaTeX sources, PDFs, PostScript, and HTML
- Structured subsets: files (54.6M rows), latex, metadata (3.15M), papertext, source, versions — auto-converted to Parquet, queryable via SQL/Dask/Polars
- License: mixed-arxiv-author-licenses (check per-paper copyright)
- A heavyweight resource for text generation and retrieval work
Related event: Entire arXiv Dumped onto Hugging Face: 3.14M Papers, 16TB, All Versions(6 posts)→
More from Infra
- GPU Programming Diary: Revisiting the Classic CUDA Matmul Optimization Worklog and MIT's Sparsity Lecture — NandoDF · 2026-09-20
- Jev-style parallel structured inference makes 350M model 63× faster, code released — helloiamleonie · 2026-09-20
- Dev builds SLO-aware inference router with Jev to pick the optimal LLM per request — ai · 2026-09-20
- Memristive Networks Learn by Reorganizing Themselves: When Material Is the Model — bravo_abad · 2026-09-20
- Crusoe signs multiyear cloud deal to run dedicated Nvidia GB300 clusters for Perplexity — Beth_Kindig · 2026-09-20
- SF Compute interviews exchange legend Rich Jaycobs on how to build a compute futures market — andriy_mulyar · 2026-09-20