Zero-data self-play pretraining: LM learns to generate its own training data from random init

inductionheads · x · 2026-09-26

Self-Play Pretraining with Zero Data starts two models from random init: a generator proposes programs for a universal Turing machine and a learner trains only on their outputs, never touching real data. Zero-shot val loss on images, text, audio and melodies decreases predictably with self-play compute, and the learner develops in-context learning. A proof-of-concept co-led by Michael Y. Li with AdityaCowsik, KfirDolev, Greg Brockman, Noah Goodman, Yoav Levine and others.

Related event: Stanford-led Research Proposes Self-Play Pretraining with Zero Data(5 posts)→

Original post →

More from Research

Research channel →