The Stack v3 arrives with 5T training tokens and 120TB of raw data
lvwerra · x · 2026-07-23
The Stack v3 ships 5T tokens for cyber-defense model training
The Stack v3 has been introduced as the dataset foundation for open code models aimed at cyber defense.
- Scale: 5T tokens ready to train, plus 120TB of raw data
- Intent: provide the training substrate for future open code models in cybersecurity and defense use cases
- Access: download link shared by the author
Related event: Hugging Face Releases The Stack v3: Largest Open-Source Code Dataset(4 posts)→
More from Safety
- Cisco releases Antares, two small models for locating known code vulnerabilities — NielsRogge · 2026-07-23
- JFPuget warns private benchmarks can leak into LLM training without no-retention terms — JFPuget · 2026-07-23
- AI safety debate turns on whether publishing process details makes open source more dangerous — BlancheMinerva · 2026-07-23
- Google’s Beyond Corp work becomes a new zero-trust model for the AI era — vijaybolina · 2026-07-23
- Press adds first-class Effects and adversarial approval for agent tool calls — RichmanRonald · 2026-07-23
- Early studies suggest spermidine may boost autophagy, mitochondria and DNA repair — Dr_Singularity · 2026-07-23