The 'Magic OOD Step' Hiding in Alignment Plans

nabla_theta · x · 2026-09-25

The author flags a sneaky pattern in alignment research: define a distribution with a simple English description that contains the thing we care about as a subset, then train on data that is by definition a different subset of that distribution. This gives off 'iid vibes' and lets people gloss over the fact that there is actually a huge out-of-distribution (OOD) shift.

The author proposes a 'no magic OOD principle': many alignment plans contain a 'magic OOD step' — train on data we can verify, then hope it generalizes to the thing we care about but cannot verify. Such steps should be actively spotted and flagged; even if they cannot be fully avoided, they should be tested via a weak-to-strong style setup that actually checks several different specific distributions.

Related event: Researcher Proposes 'No Magic OOD' Principle to Expose Flaw in Alignment Plans(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →