Falsifiable test proposed for what triggers model pain directions
EigenGender · x · 2026-09-19
- EigenGender proposes a concrete test: if a model were RL-trained to avoid randomly chosen words, those words wouldn't trigger the pain direction.
- The reasoning: pain directions should track content that is "obviously painful to an assistant persona on priors," not arbitrary negative samples.
More from AGI Musings
- Why do AI models lack personality? Safety training and attachment fears, explained — alexisgallagher · 2026-09-19
- David Patterson: If AI does every job better and cheaper, why would anyone hire you? — davidpattersonx · 2026-09-19
- AI agents are the genie: alignment failure as a modern parable of corporate greed — Michael_J_Black · 2026-09-19
- AI companies are conquering math — and exposing a discipline built on competition, not understanding — danbri · 2026-09-19
- Insurers, not regulators, will gatekeep high-risk AI evaluations, argues ex-Google policy lead — nicklaslundblad · 2026-09-19
- DeepMind's Yao Shunyu: RSI is a system problem, and post-training recipes will soon be decided by models themselves — _AndrewZhao · 2026-09-19