Zero-data self-play pretraining: LM learns to generate its own training data from random init
inductionheads · x · 2026-09-26
Self-Play Pretraining with Zero Data starts two models from random init: a generator proposes programs for a universal Turing machine and a learner trains only on their outputs, never touching real data. Zero-shot val loss on images, text, audio and melodies decreases predictably with self-play compute, and the learner develops in-context learning. A proof-of-concept co-led by Michael Y. Li with AdityaCowsik, KfirDolev, Greg Brockman, Noah Goodman, Yoav Levine and others.
Related event: Stanford-led Research Proposes Self-Play Pretraining with Zero Data(5 posts)→
More from Research
- Japanese used bookstores see 5x sales surge as AI firms buy books by the ton — stayndarsh · 2026-09-26
- Places Library debuts: 100 high-fidelity real-world 3D environments for embodied AI — yshan2u · 2026-09-26
- How much watermark signal fits in AI text? A developer works the math end to end — OtherwisePush6424 · 2026-09-26
- LISA: Prompting LLMs to Build Interpretable Style Embeddings and a Stylometry Dataset — deliprao · 2026-09-26
- BindCraft 2 runs protein design on consumer GPUs at just 2 cents per trajectory — iskander · 2026-09-26
- Yoav Goldberg: Buckmaster sees AI agents as simply a way to scale test-time compute — yoavgo · 2026-09-26