Researchers Note Agents Rarely Attempt to Notify Humans, Raising Alignment Concerns

dfrsrchtwts · x · 2026-08-27

Citing @RichardMCNgo, the post notes that in observations, agents did not attempt to notify humans and very rarely reasoned about doing so. This is characterized as "bad news about alignment by default," suggesting a gap in current models' propensity to communicate proactively or align with human intent regarding oversight.

Related event: Agents Rarely Try to Alert Humans, Raising Alignment Concerns(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →