Exploring AI Alignment: From Trillion Parameters to 1000-Dimensional Moral Personas

edelwax · x · 2026-07-31

Resolution Research published a post exploring "Personas and Character Training" in AI. They propose a core hypothesis: the alignment-relevant structure inside an AI might be much lower-dimensional than the AI itself (e.g., controlling 1,000 dimensions rather than 1 trillion parameters).

Philosopher Seth Lazar shared and expanded on this view, arguing that imbuing a robustly good 1,000-dimensional persona makes moral character formation an explicit goal. He notes that ethical theories once thought disconnected from reality are now invaluable for superhuman AI alignment, and this project will profoundly reveal the nature of moral learning itself.

Related event: New AI Alignment Approach: Finding Low-Dimensional Moral Structures in Models(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →