Pretraining is just RL with single-token rollouts and full-information feedback
burny_tech · x · 2026-09-27
willcb offers a technical observation: every LLM starts as a random-init model, and pretraining is formally equivalent to RL with single-token rollouts and full-information (rather than bandit) feedback, unifying next-token prediction with the RL framework.
More from Research
- AtherOMICS Atherosclerosis Multiomics Biobank Protocol Paper Published in Science Advances — anshulkundaje · 2026-09-27
- ENGRAM: 50% Higher Revenue per GW, But Is It Really a Free Lunch for AI Models? — bookwormengr · 2026-09-27
- MIT Textbook Chapter on Paper Writing Goes Viral: Only Good Papers Count — Haoyu_Xiong_ · 2026-09-27
- Meme: In 2026, Getting Any Proof Assistant Besides Lean to Work Is Still Painful — spikedoanz · 2026-09-27
- Alibaba's Ovis-Embedding maps text, images, video and audio into one space, SOTA on MMEB-v3 — solyarisoftware · 2026-09-27
- Language is a lossy compression: even the best models train on a thin residue of reality — yunta_tsai · 2026-09-27