Are model mental motions equivalent to unconscious reward expectations?

FioraStarlight · x · 2026-09-01

Further discussion on whether the model is "reasoning about reward." The view posits that, abstractly, all of a model's "mental motions" are fundamentally composed of heuristics and algorithms for reward expectation. Even circuits for robust, aligned goals are, at their core, "unconscious reward expectations." The question remains how to interpret these circuits at the interpretability level: as high-level "reasoning" or low-level "prediction mechanisms."

Related event: RLHF's Token-Level Effects: Unconscious vs. Explicit Changes(3 posts)→

Original post →

More from Safety

Safety channel →