Hugging Face dataset stack-v3-train is trending, with multilingual code data and arXiv linkage

HuggingFaceCode · hf · 2026-07-24

HuggingFaceCode/stack-v3-train is trending on Hugging Face as a dataset. Its metadata shows text-generation tasks, multilingual code data, a mix of crowdsourced and expert-generated sources, an ODC-BY license, and a size band between 100M and 1B examples. The dataset is also linked to arXiv:2402.19173.

Related event: Hugging Face Releases The Stack v3: 114TB Largest Open Source Code Dataset(13 posts)→

Original post →

More from Research

Research channel →