Exploration as the Third Pretraining Axis: 6.2x Data Efficiency

teortaxesTex · x · 2026-08-02

Research proposes 'exploration' as a third pretraining axis beyond parameters and data. Scaling exploration monotonically improves existing models across images, video, and language, unlocking end-to-end generation through a simple for-loop.

Gains grow with scale: data efficiency improves by 6.2x, FLOP efficiency by 4.1x, and parameter efficiency by 47%. It also achieves a near-SOTA unguided FID of 1.43 on ImageNet, allowing researchers to trade training compute for generalization.

Related event: Explorative Modeling: Third Pretraining Axis, Up to 6x Sampling Efficiency(12 posts)→

Original post →

More from Research

Research channel →