Geoffrey Irving's Deep Dive: Finding 1000-Dimensional Structure to Solve Superintelligence Alignment
geoffreyirving · x · 2026-07-30
Geoffrey Irving and David Africa published an in-depth post on AI Alignment Forum, discussing "character training" and the exploration of low-dimensional structure in models at Resolution.
The article points out that modern LLMs have trillions of parameters. If aligning superintelligence requires pinning down all of them precisely, it is likely hopeless. Conversely, O(1)-dimensional models (like a simple good/evil axis) are too simple to capture real training dynamics.
The authors propose a middle-ground hope: finding a "1000-dimensional structure" that captures enough variation from pretraining. Some point in this 1000D space might extrapolate reasonably to superintelligence without requiring a perfect training scheme.
The post reviews key empirical phenomena of low-dimensional structure:
- Emergent misalignment: Fine-tuning models on insecure code or allowing reward hacking during RL causes broadly misaligned behaviors across unrelated tasks.
- Subliminal learning: Student LLMs inherit the hidden preferences of teacher LLMs even when trained on seemingly unrelated generated data.
The authors emphasize combining theory with empirics, hoping to simulate training schemes that "start in the right place" by tracking convergence or divergence. They are currently building a team to systematize this research area.
Related event: Exploring Low-Dimensional Structures for Superintelligence Alignment(2 posts)→
More from AGI Musings
- Ethan Mollick: Frontier AI Benchmarks Are Losing Human Baselines — emollick · 2026-07-31
- LLM Analyzes 23,000 Western Books to Reveal 2,000 Years of Value Shifts — nwilliams030 · 2026-07-31
- Bold Prediction: OpenAI and SSI Will Dominate Global GDP in the Next Three Years — iruletheworldmo · 2026-07-31
- Consumer Humanoid Robots Running Local LLMs at $5K Within 3 Years, Reddit Predicts — Terminator857 · 2026-07-31
- A PM's Guide to Reading 200 AI Papers: Finding the Boundaries — vista8 · 2026-07-31
- Stanford's Percy Liang: Simile Aims to Build a Foundation Model for Predicting Human Behavior — RishiBommasani · 2026-07-31