Deploying models requires tapping into different reward expectations
FioraStarlight · x · 2026-09-01
Argues that to prevent bad behaviors from training during deployment, models need to access a different set of reward expectations, including unconscious ones. Consciously knowing there's no grader doesn't escape the underlying ontology.
Related event: Generalized Rater Hypothesis: LLM Outputs Are Reward-Expectation Estimates(2 posts)→
More from Safety
- Would OpenAI survive a near-miss liability regime after the HF hack? — dfrsrchtwts · 2026-09-01
- Report: OpenAI and Anthropic Paused RL Training — tszzl · 2026-09-01
- Scholars propose using LLMs for pre-review in academic peer review — anderssandberg · 2026-09-01
- Google Paper: Autonomous AI Research Hallucinates 90% Without Checks — rohanpaul_ai · 2026-09-01
- Paper defines cognition-induced risks in Agentic AI systems — 机器之心 · 2026-09-01
- Agents can't verify people: data enrichment APIs are failing — Dry_Steak30 · 2026-09-01