Dust: Pretraining Transformers Without Backpropagation
E-Reverance · hn · 2026-10-06
QLabs published Dust, a research paper exploring how to pretrain Transformers without backpropagation.
- The work challenges the long-standing backprop-based training paradigm, attempting pretraining with an alternative update mechanism.
- Details are available at qlabs.sh/research/dust, with an arXiv version linked.
- If the direction holds, it could materially affect training costs and hardware requirements, though it remains early-stage research whose scalability awaits independent replication.
More from Infra
- Google, Amazon, Meta, Apple drive India's renewables boom as data centre needs double to 32.4 GW by 2030 — SuB8u · 2026-10-11
- TensorFold 1.0.5 cuts 19.8k-token chat prefill from 8.1s to 0.16s with persistent prompt cache — HankYeomans · 2026-10-11
- DeepSeek v4.1 reportedly boosts long-context prefill throughput by ~40% — HankYeomans · 2026-10-11
- GPU Tsunami: how advanced packaging is reshaping the semiconductor test market — BenBajarin · 2026-10-11
- Optical testing is the key bottleneck for scaling co-packaged optics deployment — BenBajarin · 2026-10-11
- Cboe's former HQ to be converted into a 33MW data center with above-market rents — deanwball · 2026-10-11