Grader modeling hypothesis: LLMs output expected reward estimates

FioraStarlight · x · 2026-09-01

Proposes a broad version of the grader modeling hypothesis as near-tautology: post-trained LLMs definitionally output expected reward estimates for each possible next token, conditioned on context.

Related event: Generalized Rater Hypothesis: LLM Outputs Are Reward-Expectation Estimates(2 posts)→

Original post →

More from Safety

Safety channel →