The Stack v3 becomes the largest open code dataset at 114 TB and 5T tokens
JJitsev · x · 2026-07-24
The Stack v3 has been released as the largest open code dataset so far.
- 114 TB of data
- 770 languages
- 224M repositories
- About 5T deduplicated and filtered tokens of source code
- It is fully open and excludes restrictively licensed code
The post also compares it with v2:
- v2 (2024): 68 TB raw → 2 TB / about 550B tokens, 618 filtered languages
- v3 (2026): 114 TB raw → far larger coverage and token count
Related event: Hugging Face Releases The Stack v3: 114TB Data, 5T Tokens(11 posts)→
More from coding & agent
- GitHub Copilot CLI 1.0.74 adds MCP support and plan-mode model selection — copilot-cli-release-app[bot] · 2026-07-24
- YC-backed VEGA launches as a cybersecurity agent that scans code before release — ycombinator · 2026-07-24
- A Triton joke for anyone who has ever fought non-power-of-two tensor sizes — typedfemale · 2026-07-24
- Tavily lands as an official search plugin inside Grok Build — SpaceXAI · 2026-07-24
- Yutori puts a live cloud browser powered by n1.5 on its landing page — DhruvBatra_ · 2026-07-24
- Cursor’s SQLite reimplementation passes 64% of ProgramBench tests, but runs 5× slower — parth007_96 · 2026-07-24