The Stack v3 becomes the largest open code dataset at 114 TB and 5T tokens
JJitsev · x · 2026-07-24
The Stack v3 has been released as the largest open code dataset so far.
- 114 TB of data
- 770 languages
- 224M repositories
- About 5T deduplicated and filtered tokens of source code
- It is fully open and excludes restrictively licensed code
The post also compares it with v2:
- v2 (2024): 68 TB raw → 2 TB / about 550B tokens, 618 filtered languages
- v3 (2026): 114 TB raw → far larger coverage and token count
Related event: Hugging Face Releases The Stack v3: 114TB Largest Open Source Code Dataset(13 posts)→
More from coding & agent
- MathModelAgent gains traction: auto-solves math modeling and writes a submission-ready paper — jihe520 · 2026-09-11
- alphaXiv open-sources OpenResearch to run parallel research agents with any model — alphaXiv · 2026-09-11
- DeskcommCRM: open-source AI sales CRM with native agents and WhatsApp hits 1k stars — melgarafael · 2026-09-11
- hyperresearch: agent-driven knowledge base that turns web research into a searchable wiki — jordan-gibbs · 2026-09-11
- Forter's 13 lessons from its agent sprint: skip custom RAG, lean on mature enterprise search — bibryam · 2026-09-11
- Two real 'company brains' opened up live: Gorgias' in-house Cortex vs Slite — femke_plantinga · 2026-09-11