Researchers release 44B synthetic tokens for higher-quality pretraining data
vanstriendaniel · x · 2026-07-21
Researchers release 44B synthetic tokens from CoT-guided rewriting
The post announces a release of 44B synthetic tokens generated through CoT-guided rewriting.
Main claim
- The authors say the resulting data offers higher-quality pretraining material than the average human-written web text.
- The release includes both the dataset and the paper.
- The paper has been accepted at COLM 2025.
Why it matters
- It is aimed at improving pretraining data quality, not just scale.
- The work suggests synthetic text can be filtered and rewritten in a way that may outperform raw web data for model training.
More from Research
- Project APE launches CRED to test whether LLMs can verify research errors — soumitrashukla9 · 2026-07-22
- Project APE finds verifier reliability drops when papers contain multiple errors — soumitrashukla9 · 2026-07-22
- Project APE says verifier costs fell about 90x in a year as Chinese open models lead — soumitrashukla9 · 2026-07-22
- OpenAI-linked paper says capability RL can make models more reward-seeking — MariusHobbhahn · 2026-07-22
- Project APE builds its verifier benchmark from 100 AI-written papers with injected errors — soumitrashukla9 · 2026-07-22
- Paper proposes a CRED taxonomy and benchmark to measure research-error detectors — soumitrashukla9 · 2026-07-22