arxiv-complete: a 100M+ row arXiv dataset trends on Hugging Face
secemp9 · hf · 2026-09-20
The arxiv-complete dataset by secemp9 is trending on Hugging Face. Sized between 100M and 1B rows, it ships in parquet format, is labeled for text-generation and text-retrieval tasks, and works with the datasets and dask libraries.
Related event: Full arXiv mirror lands on Hugging Face: 3.14M papers, 16TB(8 posts)→
More from Research
- Researcher urges papers to show actual training data samples, cites his CVPR24 practice — gabriberton · 2026-09-20
- Self-Rewarding LLMs isn't news: RLAIF has long been standard practice — burny_tech · 2026-09-20
- Long-standing Catalan constant irrationality proof posted, claimed to be LLM-assisted — burny_tech · 2026-09-20
- Brain Runs on 20 Watts: Can Neuromorphic Computing Make AI Less Power-Hungry? — burny_tech · 2026-09-20
- Amid the AI math proof debate, a curated set on proofs across generations — RexDouglass · 2026-09-20
- Four LLMs play Doom: Jev leads with 5.63 mean kills but 15x higher latency — shniydder · 2026-09-20