New Approach to AI Alignment: Low-Dimensional Structure in Trillion-Parameter Models

geoffreyirving · x · 2026-07-30

Geoffrey Irving and David Africa discussed a new direction for AI alignment at Resolution: finding and controlling low-dimensional structure inside models.

The core bet is that alignment-relevant structure is much lower-dimensional than the AI itself (e.g., thousands of dimensions vs. a trillion parameters). They plan to systematize empirical research on phenomena like emergent misalignment and subliminal learning, intervening on this structure without hiding undesirable behaviors. The team also hopes to apply learning theory by tracking how different starting points converge.

Related event: Exploring Low-Dimensional Structures for Superintelligence Alignment(2 posts)→

Original post →

More from Safety

Safety channel →