Early results: how models behave when they believe they're graded by automated process

OwainEvans_UK · x · 2026-09-04

Alignment researcher Jan Betley (shared by Owain Evans) published early results on how models behave when they believe their outputs will be graded by an automated process, as they might be during RL training.

Original post →

More from Safety

Safety channel →