FinePhrase (COLM Oral): 1T-Token Study Finds Structured Synthetic Data Beats Curated Web, Cuts Costs 30x
edwardbeeching · x · 2026-10-06
- FinePhrase, accepted as an Oral Spotlight at COLM 2026, is the first systematic study of synthetic pretraining data design (rephrasing strategy, generator model, source data), based on controlled experiments generating over 1 trillion tokens.
- Findings: structured output formats (tables, math problems, FAQs, tutorials) consistently beat curated web baselines and prior synthetic methods; generator models beyond 1B parameters add no benefit; source data selection matters substantially.
- Output: a 486B-token open dataset of rephrased web text that outperforms all synthetic baselines while cutting generation costs by up to 30x; dataset, prompts and framework are open-sourced. Talk on Wednesday at COLM in SF.
More from Research
- Neuroscience Debate: Why Engram Research Lags Behind Staining and Patch-Clamp Breakthroughs — ShahabBakht · 2026-10-06
- CV4EO workshop returns to WACV with paper deadline Oct 14, 2026 — abby621 · 2026-10-06
- H-JEPA Follow-Up: SIGReg Keeps Latent Spaces Well-Conditioned Across All Levels — randall_balestr · 2026-10-06
- Perspective paper sparks debate: causal experiments can't answer causal questions without a model of the neural code — ShahabBakht · 2026-10-06
- Ben Goertzel unveils 'Distinction-Calculus', new math for forking, merging, self-modifying agent minds — bengoertzel · 2026-10-06
- Looped Transformers Enable Depth Extrapolation, Fueling Claude Mythos Architecture Speculation — heghbalz · 2026-10-06