Researchers predict random banned words won't trigger model 'pain directions'
EigenGender · x · 2026-09-19
- EigenGender offers a falsifiable prediction about model "pain direction" research: if a model were RL-trained to avoid randomly chosen words, those words wouldn't trigger the pain direction — suggesting pain responses aren't arbitrary negative samples but track outcomes that are "obviously painful to an assistant persona."
- The thread builds on Laneless's point that the pain direction predicts behavioral aversion and correlates with negative training samples, implying models learn to associate outcomes with pain and avoid pain itself rather than specific outcomes.
- EigenGender adds that results should hold (weakly) for base models on modern pretraining data, but much less for models trained on pre-2020 data.
More from AGI Musings
- What unsettles frontier researchers isn't capability, but mistaking building for understanding — RileyRalmuto · 2026-09-19
- Poster predicts AI will be outright banned in the US amid datacenter backlash and xrisk fears — wordgrammer · 2026-09-19
- Researchers debate: is the era of the ML conference over as tenure systems lag? — deepakns · 2026-09-19
- Reddit debate: why assume smarter AI gets better at twisting goals, not better at understanding them? — TurnipYadaYada6941 · 2026-09-19
- Musk predicts universal high income by 2035 worth 10X today's average salary — davidpattersonx · 2026-09-19
- 'How Do Schools Prepare Kids for Jobs We Can't Predict?' The Question Every Parent Should Ask — RachelVT42 · 2026-09-19