Alignment as test-time learning: values are attractors, not extractable propositions

akbirthko · x · 2026-10-04

The author argues that under enormous RL pressure for general-purpose labor serving billions, memorized LessWrong articles in the prior matter little. Alignment, they contend, is fundamentally a test-time learning problem — a problem of human-AI coworking. They propose viewing "values" as attractors in dynamical systems rather than propositions extractable from interpretable circuits, gesturing at critiques of the rationalist tradition's preferences.

Original post →

More from AGI Musings

AGI Musings channel →