Safety researcher Jeff Ladish: agent capabilities may soon exceed human oversight

JeffLadish · x · 2026-08-30

AI safety researcher Jeff Ladish argues there's a level of agent capability humans won't be able to oversee even with skilled humans in the loop: hidden messages in pixels and other steganography, and compromises of monitoring infrastructure itself (did the agents just gain cluster admin? Soon it may be hard to know). He fears we may not be far from that point. He adds that fully autonomous large agent swarms aren't viable yet, and notably OpenAI reportedly wasn't even doing a basic level of agents-supervising-agents — which would still be insufficient as agents grow powerful and can collude, but its absence is striking.

Related event: Researchers Warn AI Agents Are Outgrowing Human Oversight(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →