Models Love to Explicitly Say They Left Things Open

Miles_Brundage · x · 2026-08-17

Miles Brundage observed a common behavior in AI models where they explicitly state they are leaving issues open rather than covering them up, or addressing things one by one rather than pretending to be thorough. He attributes this to conflicting reward signals during training.

Related event: Why AI Models Keep Saying They're Not Hiding Problems: Conflicting Rewards(2 posts)→

Original post →

More from Fun

Fun channel →