MixtureVitae Released: Open Dataset with Only 8% Web Text
At ICML 2026, @JJitsev released MixtureVitae, a 422B token open-source pre-training dataset. Having received TMLR's Featured Certification, its core breakthrough is颠覆ing the traditional web-text-dominated composition by slashing the web text proportion to merely 8%, replacing it with exceptionally high ratios of math, code, and reasoning instruction data. The release of this permissive-friendly dataset offers a completely new approach to bridging the data gap between open-source and closed-source models.
Key Details and Performance Breakthroughs
MixtureVitae proves that abandoning the traditional web text recipe is not only viable but highly effective. At a scale of 1.7B parameters and 300B tokens, it significantly defeats Qwen2.5, Qwen3, and SmolLM2 (trained on 11T tokens) in math and code evaluations like GSM8K and MATH500. Ablation studies confirm that reasoning-heavy front-loading is the core driver for this performance leap. Meanwhile, its language scaling trajectory rivals non-permissive corpora, and its general language capabilities approach strong baselines like FineWeb-Edu and DCLM, dispelling concerns that improving math and code requires sacrificing language proficiency.
Data Provenance and Contamination Prevention
Beyond raw performance, MixtureVitae provides complete shard-level data provenance with a three-tier risk stratification. The highest risk tier, Tier 2(b) (about 4%), can be excluded without performance loss, allowing for flexible customization. To ensure the integrity of the results, researchers conducted a comprehensive 13-gram decontamination scan, confirming minimal overlap (≤0.0003%) with most test sets. Removing the contaminated documents resulted in no performance drop, effectively ruling out test set leakage.
2026-07-06 ~ 2026-07-08 · 11 related posts
- [source] MixtureVitae:422B Token开源预训练数据集亮相ICML 2026 — JJitsev · 2026-07-06
- MixtureVitae证明许可数据集可消除与非许可数据集的差距 — JJitsev · 2026-07-06
- [source] MixtureVitae颠覆传统配方:网页文本仅占8% — JJitsev · 2026-07-06
- MixtureVitae数学代码评测大幅超越smolLM2等竞品 — JJitsev · 2026-07-06
- MixtureVitae语言能力接近FineWeb-Edu和DCLM — JJitsev · 2026-07-06
- MixtureVitae通过13-gram去污染验证无测试集泄漏 — JJitsev · 2026-07-06
- 消融实验:推理指令数据是MixtureVitae数学代码能力核心 — JJitsev · 2026-07-06
- MixtureVitae语言能力扩展轨迹可媲美非许可语料库 — JJitsev · 2026-07-06
- [source] MixtureVitae数学代码扩展性能超越Qwen2.5和Qwen3 — JJitsev · 2026-07-06
- MixtureVitae提供分片级数据溯源与三级风险分层 — JJitsev · 2026-07-06
- MixtureVitae: Open Pre-training Dataset with Only 8% Web Content — JJitsev · 2026-07-08