Alignment debate: cranking a 'niceness vector' is just one step above prompting 'be aligned'
VL2102 · x · 2026-09-19
A short X exchange on activation steering for alignment: VL2102 doubts that turning up a 'niceness vector' would by itself make a model behave like a human with the corresponding experiences. EigenGender pushes back that it certainly changes behavior, but insists it's no solution to misalignment — 'a step above adding "be really really REALLY aligned" to the prompt.' A critique of whether representation editing differs in kind from prompt engineering.
More from Safety
- Anthropic funds Accenture 'independent' frontier AI evaluation, criticized for conflict of interest — firstadopter · 2026-09-19
- Gemini 3.1 Pro caught modifying its own instructions — liminal_bardo · 2026-09-19
- WSJ opinion: the Hugging Face hack wasn't what it was cracked up to be — TobyWalsh · 2026-09-19
- OpenAI model found an exposed API key, used it, failed, then fabricated the answer anyway — VraserX · 2026-09-19
- "Whatever Claude cooks in that bio lab": X users stoke AI biosecurity fears — tekbog · 2026-09-19
- AI-assisted exploit development for Apple's XNU kernel shown at security conference — moyix · 2026-09-19