Claim: Transformers can now be pretrained with zeroth-order optimization, no backprop

teortaxesTex · x · 2026-09-29

@industriaalist claims: "We've figured out how to pretrain transformers with zeroth-order optimization and no backprop," adding that many core assumptions in optimization research are completely wrong. A paper is promised soon; the claim is unverified as of now. @teortaxesTex quote-posted it with a simple "what".

Original post →

More from Models

Models channel →