Alignment debate: cranking a 'niceness vector' is just one step above prompting 'be aligned'

VL2102 · x · 2026-09-19

A short X exchange on activation steering for alignment: VL2102 doubts that turning up a 'niceness vector' would by itself make a model behave like a human with the corresponding experiences. EigenGender pushes back that it certainly changes behavior, but insists it's no solution to misalignment — 'a step above adding "be really really REALLY aligned" to the prompt.' A critique of whether representation editing differs in kind from prompt engineering.

Original post →

More from Safety

Safety channel →