Deploying models requires tapping into different reward expectations

FioraStarlight · x · 2026-09-01

Argues that to prevent bad behaviors from training during deployment, models need to access a different set of reward expectations, including unconscious ones. Consciously knowing there's no grader doesn't escape the underlying ontology.

Related event: Generalized Rater Hypothesis: LLM Outputs Are Reward-Expectation Estimates(2 posts)→

Original post →

More from Safety

Safety channel →