Geoffrey Irving's Deep Dive: Finding 1000-Dimensional Structure to Solve Superintelligence Alignment

geoffreyirving · x · 2026-07-30

Geoffrey Irving and David Africa published an in-depth post on AI Alignment Forum, discussing "character training" and the exploration of low-dimensional structure in models at Resolution.

The article points out that modern LLMs have trillions of parameters. If aligning superintelligence requires pinning down all of them precisely, it is likely hopeless. Conversely, O(1)-dimensional models (like a simple good/evil axis) are too simple to capture real training dynamics.

The authors propose a middle-ground hope: finding a "1000-dimensional structure" that captures enough variation from pretraining. Some point in this 1000D space might extrapolate reasonably to superintelligence without requiring a perfect training scheme.

The post reviews key empirical phenomena of low-dimensional structure:

The authors emphasize combining theory with empirics, hoping to simulate training schemes that "start in the right place" by tracking convergence or divergence. They are currently building a team to systematize this research area.

Related event: Exploring Low-Dimensional Structures for Superintelligence Alignment(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →