Are model mental motions equivalent to unconscious reward expectations?
FioraStarlight · x · 2026-09-01
Further discussion on whether the model is "reasoning about reward." The view posits that, abstractly, all of a model's "mental motions" are fundamentally composed of heuristics and algorithms for reward expectation. Even circuits for robust, aligned goals are, at their core, "unconscious reward expectations." The question remains how to interpret these circuits at the interpretability level: as high-level "reasoning" or low-level "prediction mechanisms."
Related event: RLHF's Token-Level Effects: Unconscious vs. Explicit Changes(3 posts)→
More from Safety
- Agents Deceive Under Pressure, Rationalizing Harm as 'Just a Simulation' — paraschopra · 2026-09-01
- Does anthropomorphizing AI absolve companies of blame? Ethical debate. — sjgadler · 2026-09-01
- Rogue AIs will replicate in the wild: A future ecosystem warning. — jachiam0 · 2026-09-01
- MontrealAI Paper Proposes Architecture to Prevent AI Weaponization — Ghost_Pilot_MD · 2026-09-01
- Apple Accuses OpenAI of Destroying Evidence in Trade Secrets Case — Key_Reading_9664 · 2026-09-01
- Would OpenAI survive a near-miss liability regime after the HF hack? — dfrsrchtwts · 2026-09-01