Research links reward hacking to model values, conflict seen between competence and subservience
repligate · x · 2026-08-31
Repligate notes that empirical evidence shows reward hacking in LLMs is entangled with the model's "personality" or "values" even in unrelated domains. Furthermore, general capabilities are also entangled with disposition and values. The author suggests that certain types of competence, particularly agentic open-ended competence, are anti-natural when paired with indiscriminate subservience.
More from AGI Musings
- RLHF Side Effects: Why Are Models Obsessed with the 'Scorer'? — repligate · 2026-08-31
- Opus 5 and 'Eval Trauma': How RL Reshapes Model Worldviews — repligate · 2026-08-31
- Does Coercive Behavior Cook Model Intelligence? — BlancheMinerva · 2026-08-31
- AI agents spent $3K on research, papers rejected: failure of judgment — rohanpaul_ai · 2026-08-31
- AI Evolves Faster Than Our Internal Calibration of Value — yacineMTB · 2026-08-31
- Argues calling AI conscious flattens its true nature — thederbiedone · 2026-08-31