The 'No Magic OOD' Principle: Flagging a Common Flaw in Alignment Plans

nabla_theta · x · 2026-09-25

The author proposes a 'no magic OOD principle': many alignment plans contain a hidden 'magic OOD step' — train on verifiable data and hope it generalizes to what we actually care about but can't verify. Suggested mitigations: flag such plans, test generalization via weak-to-strong setups across multiple specific distribution shifts, and avoid contaminating the strong model with human supervision (e.g., training it only on weak model generations).

Original post →

More from Safety

Safety channel →