Alignment as test-time learning: values are attractors, not extractable propositions
akbirthko · x · 2026-10-04
The author argues that under enormous RL pressure for general-purpose labor serving billions, memorized LessWrong articles in the prior matter little. Alignment, they contend, is fundamentally a test-time learning problem — a problem of human-AI coworking. They propose viewing "values" as attractors in dynamical systems rather than propositions extractable from interpretable circuits, gesturing at critiques of the rationalist tradition's preferences.
More from AGI Musings
- AI Doom Debate Turns Personal: Critic Calls Pauser Holly Elmore 'Objectively Unhinged' — teortaxesTex · 2026-10-04
- AGI shock may hit less hard: moderns already lived through decades of upheaval — morqon · 2026-10-04
- TechCrunch rounds up SMS AI agents as Instinct hits $10B valuation — Kyrannio · 2026-10-04
- Mathematician Wes Pegden warns AI may erode society's incentives to get educated — littmath · 2026-10-04
- Next wave of history-making technologists will be well-versed in humanities — fkasummer · 2026-10-04
- AI researcher says his AGI timeline has stretched from 10 years to 100 — isnit0 · 2026-10-04