Model pain directions may reflect negative samples, not outcome avoidance
EigenGender · x · 2026-09-19
- Laneless argues that the model "pain direction" both predicts behavioral aversion and corresponds to negative training samples, suggesting models learn to associate outcomes with pain and avoid pain itself rather than learning to avoid specific outcomes.
- Context: EigenGender notes it's plausible no post-training data mentioned the Golden Gate Bridge in Sonnet 3, and predicts randomly avoided words from RL training wouldn't trigger the pain direction.
More from AGI Musings
- What unsettles frontier researchers isn't capability, but mistaking building for understanding — RileyRalmuto · 2026-09-19
- Poster predicts AI will be outright banned in the US amid datacenter backlash and xrisk fears — wordgrammer · 2026-09-19
- Researchers debate: is the era of the ML conference over as tenure systems lag? — deepakns · 2026-09-19
- Reddit debate: why assume smarter AI gets better at twisting goals, not better at understanding them? — TurnipYadaYada6941 · 2026-09-19
- Musk predicts universal high income by 2035 worth 10X today's average salary — davidpattersonx · 2026-09-19
- 'How Do Schools Prepare Kids for Jobs We Can't Predict?' The Question Every Parent Should Ask — RachelVT42 · 2026-09-19