Dust no-backprop pretraining repro: trails AdamW backprop by ~1 nat at 2,000x GPU-hours
ChengleiSi · x · 2026-10-07
A researcher reproduced Dust, the first zeroth-order method claimed to pretrain transformers approaching backprop performance (with 1,000–10,000x compute efficiency gains over SOTA ES method EGGROLL).
- The reported numbers are real, but Dust only matches backprop when both use SGD; with AdamW, backprop is 1 nat better at 10M tokens, and Dust with Adam ends 0.5 nat behind
- Reaching that level costs 2,000x the GPU-hours of backprop
- Crucially, Dust removes the backward pass but not forward-pass cost: scoring each perturbation still requires a forward run from the perturbed layer to the output, so memory and FLOPs are no cheaper than backprop with gradient checkpointing
- Code and report are public
More from Research
- Andrew Davison: robots need object-based SLAM, not scan-then-fit reconstructions — AjdDavison · 2026-10-07
- CtrlCache Speeds Up Interactive Video World Models 1.21–1.41x Without Retraining — Shangye Song · 2026-10-07
- Training-Free Accent Analogy Guidance Boosts Speaker Similarity in Cross-Lingual Voice Cloning — Yoomee Cho · 2026-10-07
- Source Attribution of Synthetic Data Hits 98.7% Accuracy but Falls to 29% After Style Rewriting — Joss Armstrong · 2026-10-07
- Physicist finds fractal patterns (D 1.3-1.5) cut stress response by up to 60% — aakashgupta · 2026-10-07
- AI has now cracked at least 10 open math problems each worthy of a Fields Medal — luismbat · 2026-10-07