Eight months of GEN-1.5 pretraining sparks essay on scaling as universal methodology

CatAstro_Piyush · x · 2026-09-07

LLM pretraining researcher Jiaxuan Zou argues, starting from GEN-1.5's eight-month loss curve (validation action-prediction error across three training stages), that pretraining and scaling form a universal methodology: data supplies experience, compute supplies computation, and training encodes structure from that experience into weights — shared by language models, world models, and embodied AI.

Citing Yang Zhilin's criteria, the essay says success requires a general framework (architecture, data representation, objective, optimizer) and a scalable learning process. Evidence: Dyna-2 scales human video pretraining to 1 million hours, improving prediction on unseen robot data, with real-world task performance improving alongside pretraining data scale under the same post-training setup.

Original post →

More from Embodied

Embodied channel →