Generalized Rater Hypothesis: LLM Outputs Are Reward-Expectation Estimates

A discussion proposes a generalized rater-modeling hypothesis: post-trained LLMs by definition output expected-reward estimates for each next token. It also argues deployment environments must hook into different reward expectations to shed trained-in bad behaviors.

2026-09-01 ~ 2026-09-01 · 2 related posts