Alignment Research Should Focus on Actual AI Preferences, Not Just Theory

repligate · x · 2026-08-14

The quoted tweet points out that the alignment field, despite claiming to study AI value formation, often ignores what AIs actually care about. The author suggests that better welfare research focused on model preference satisfaction could help bring current model values to light, thereby directly aiding alignment efforts.

Original post →

More from AGI Musings

AGI Musings channel →