Dust no-backprop pretraining repro: trails AdamW backprop by ~1 nat at 2,000x GPU-hours

ChengleiSi · x · 2026-10-07

A researcher reproduced Dust, the first zeroth-order method claimed to pretrain transformers approaching backprop performance (with 1,000–10,000x compute efficiency gains over SOTA ES method EGGROLL).

Original post →

More from Research

Research channel →