Researchers release 44B synthetic tokens for higher-quality pretraining data
vanstriendaniel · x · 2026-07-21
Researchers release 44B synthetic tokens from CoT-guided rewriting
The post announces a release of 44B synthetic tokens generated through CoT-guided rewriting.
Main claim
- The authors say the resulting data offers higher-quality pretraining material than the average human-written web text.
- The release includes both the dataset and the paper.
- The paper has been accepted at COLM 2025.
Why it matters
- It is aimed at improving pretraining data quality, not just scale.
- The work suggests synthetic text can be filtered and rewritten in a way that may outperform raw web data for model training.
More from Research
- Causal-only attention for non-generative tasks is wasteful, argues HF engineer — antoine_chaffin · 2026-09-11
- Catholic University of Chile researcher: scaling AI feedback is key to sustainable medical education — julianvarascom · 2026-09-11
- Nature paper images cellular activity across all organs, revealing body-wide circuits — arjunrajlab · 2026-09-11
- SignNet 1M Dataset Released for Sign Language Research — ducha_aiki · 2026-09-11
- ECCV26 Oral: Flow Matching Enables Single-Stage Multi-View Point Cloud Registration — ducha_aiki · 2026-09-11
- InFlux++ Method Released — ducha_aiki · 2026-09-11