'Posttraining is translation' partially retracted: verifiers can pull mass beyond the base model
Liu_eroteme · x · 2026-09-02
In a technical thread, the author partially retracts their earlier claim that 'posttraining is translation.' While posttraining resembles translation over the bulk of the distribution, at the frontier the verifier can pull probability mass into sequences the base model would never produce — which explains why AlphaGo could discover novel strategies: with a single verifier axis, the basin gets stretched arbitrarily far. The author also contrasts pretraining as policy-gradient optimization on a maximally dense distributional reward (with a global basin and nonzero gradient from any initialization) against bandit feedback, where gradients exist only where the policy already places mass.
More from Research
- New paper shows Fiat-Shamir transformation breaks soundness for program-generated R1CS proof systems — jedisct1 · 2026-09-02
- TwinDEX: wearable co-designed grippers let a dexterous robot finish 20+ step chemistry lab with zero on-robot training — chris_j_paxton · 2026-09-02
- Cambridge team combines ML with conventional solvers for AC optimal power flow — lawrennd · 2026-09-02
- Sander Dieleman's 'Perspectives on Diffusion' resurfaces in flow matching debate — psuraj28 · 2026-09-02
- NUS researcher Xue Fuzhao: loop transformer ideas are cheap, careful execution is the AGI route — XueFz · 2026-09-02
- AITHYRA opens fully funded PhD call for AI × biomedical research in Vienna — LucaAmb · 2026-09-02