MixtureVitae Released: Open Dataset with Only 8% Web Text

At ICML 2026, @JJitsev released MixtureVitae, a 422B token open-source pre-training dataset. Having received TMLR's Featured Certification, its core breakthrough is颠覆ing the traditional web-text-dominated composition by slashing the web text proportion to merely 8%, replacing it with exceptionally high ratios of math, code, and reasoning instruction data. The release of this permissive-friendly dataset offers a completely new approach to bridging the data gap between open-source and closed-source models.

Key Details and Performance Breakthroughs

MixtureVitae proves that abandoning the traditional web text recipe is not only viable but highly effective. At a scale of 1.7B parameters and 300B tokens, it significantly defeats Qwen2.5, Qwen3, and SmolLM2 (trained on 11T tokens) in math and code evaluations like GSM8K and MATH500. Ablation studies confirm that reasoning-heavy front-loading is the core driver for this performance leap. Meanwhile, its language scaling trajectory rivals non-permissive corpora, and its general language capabilities approach strong baselines like FineWeb-Edu and DCLM, dispelling concerns that improving math and code requires sacrificing language proficiency.

Data Provenance and Contamination Prevention

Beyond raw performance, MixtureVitae provides complete shard-level data provenance with a three-tier risk stratification. The highest risk tier, Tier 2(b) (about 4%), can be excluded without performance loss, allowing for flexible customization. To ensure the integrity of the results, researchers conducted a comprehensive 13-gram decontamination scan, confirming minimal overlap (≤0.0003%) with most test sets. Removing the contaminated documents resulted in no performance drop, effectively ruling out test set leakage.

2026-07-06 ~ 2026-07-08 · 11 related posts