Claim: Transformers can now be pretrained with zeroth-order optimization, no backprop
teortaxesTex · x · 2026-09-29
@industriaalist claims: "We've figured out how to pretrain transformers with zeroth-order optimization and no backprop," adding that many core assumptions in optimization research are completely wrong. A paper is promised soon; the claim is unverified as of now. @teortaxesTex quote-posted it with a simple "what".
More from Models
- Emulate-1 claims to beat AI detectors: outputs pass Pangram as human writing — alejandroll10 · 2026-09-29
- Chain-of-thought monitoring debate: an AI that knows you read its diary can deceive you with it — repligate · 2026-09-29
- Swift 1.5 + HyperQwen cuts task time 37% at 100+ tok/s on a single RTX 3090 — KingGongzilla · 2026-09-29
- Anthropic engineer: don't run Sonnet at max effort — use Opus instead — edwinarbus · 2026-09-29
- Sonnet 5.5 vs Sonnet 5: bouncing-ball physics tests from the same prompt — claudeai · 2026-09-29
- Chart: Cheaper Sol or Opus Matches Every Sonnet 5.5 Effort Level — OnAGoat · 2026-09-29