Hugging Face releases The Stack v3, a 114 TB open code corpus with train and full buckets
Nunki08 · reddit · 2026-07-24
Hugging Face released The Stack v3, describing it as the largest open code dataset yet.
- stack-v3-train is near-deduplicated, quality-filtered, PII-redacted, and ready to load with loaddataset.
- stack-v3-full exposes the full 114 TB corpus as an HF Storage Bucket, keeping duplicates and cluster IDs so users can build their own dedup/filter/mixing pipelines.
- The post highlights both the curated training split and the raw corpus as separate entry points for researchers and builders.
More from Research
- FinanceComplexQA adds a 2,026-task benchmark for agentic reasoning on financial docs — Beihang · 2026-07-24
- Microsoft Research’s ReOPD reuses teacher prefixes to distill multi-turn agents offline — MicrosoftResearch · 2026-07-24
- Dive into LLMs tutorial repo jumps to 44,837 GitHub stars — Lordog · 2026-07-24
- Bifrost says its real2sim pipeline can rebuild a site video into a simulation-ready 3D world in 30 minutes — jnack · 2026-07-24
- Meta’s GAMUT benchmark scores long answers on missing facts, and the best model gets 58.7% — rohanpaul_ai · 2026-07-24
- Masked Visual Actions turns 15 hours of robot video into a zero-shot world model — jbhuang0604 · 2026-07-24