'Posttraining is translation' partially retracted: verifiers can pull mass beyond the base model

Liu_eroteme · x · 2026-09-02

In a technical thread, the author partially retracts their earlier claim that 'posttraining is translation.' While posttraining resembles translation over the bulk of the distribution, at the frontier the verifier can pull probability mass into sequences the base model would never produce — which explains why AlphaGo could discover novel strategies: with a single verifier axis, the basin gets stretched arbitrarily far. The author also contrasts pretraining as policy-gradient optimization on a maximally dense distributional reward (with a global basin and nonzero gradient from any initialization) against bandit feedback, where gradients exist only where the policy already places mass.

Original post →

More from Research

Research channel →