Generalized Rater Hypothesis: LLM Outputs Are Reward-Expectation Estimates
A discussion proposes a generalized rater-modeling hypothesis: post-trained LLMs by definition output expected-reward estimates for each next token. It also argues deployment environments must hook into different reward expectations to shed trained-in bad behaviors.
2026-09-01 ~ 2026-09-01 · 2 related posts
- Grader modeling hypothesis: LLMs output expected reward estimates — FioraStarlight · 2026-09-01
- Deploying models requires tapping into different reward expectations — FioraStarlight · 2026-09-01