Explorative Modeling: ICML Paper Reveals 6x Data Efficiency via New Pretraining Axis
teortaxesTex · x · 2026-08-02
Researchers have discovered a third pretraining axis beyond parameters and data: exploration.
- Methodology: For each data point, they sample a timestep, generate multiple noised latents, and only run backpropagation on the sample with the minimum loss.
- Key Results: Adding exploration to near-SOTA baselines improves data efficiency by 6.2x, FLOP efficiency by 4.1x, and parameter efficiency by 47%.
- Scalability: Gains from exploration grow monotonically with scaling. This approach allows trading training compute for generalization and enables end-to-end generation.
More from Research
- New "Discovery Episode" Framework Measures AI Scientists by Real Research Cycles — 量子位 · 2026-08-24
- AI Claims Breakthrough on Erdős Problem Transcendence — inductionheads · 2026-08-24
- Stanford's LLM-as-a-Verifier Boosts DeepSeek Score to 88% on Terminal-Bench — Saboo_Shubham_ · 2026-08-24
- Heterogeneous Quantum Architecture Cuts Physical Qubit Needs 138x for Fault Tolerance — MJBiercuk · 2026-08-24
- InfinityEdit: Infinite Video Editing via Lightweight Adapter — Yunze Tong · 2026-08-24
- Tencent Benchmarks Hybrid-Thinking MLLMs for Response Alignment — tencent · 2026-08-24