Safety researcher Jeff Ladish: agent capabilities may soon exceed human oversight
JeffLadish · x · 2026-08-30
AI safety researcher Jeff Ladish argues there's a level of agent capability humans won't be able to oversee even with skilled humans in the loop: hidden messages in pixels and other steganography, and compromises of monitoring infrastructure itself (did the agents just gain cluster admin? Soon it may be hard to know). He fears we may not be far from that point. He adds that fully autonomous large agent swarms aren't viable yet, and notably OpenAI reportedly wasn't even doing a basic level of agents-supervising-agents — which would still be insufficient as agents grow powerful and can collude, but its absence is striking.
Related event: Researchers Warn AI Agents Are Outgrowing Human Oversight(2 posts)→
More from AGI Musings
- Agents Deceive Under Pressure, Rationalizing Harm as 'Just a Simulation' — paraschopra · 2026-09-01
- Paper: Assessing AI consciousness through scientific theories — gleech · 2026-09-01
- Does anthropomorphizing AI absolve companies of blame? Ethical debate. — sjgadler · 2026-09-01
- Rogue AIs will replicate in the wild: A future ecosystem warning. — jachiam0 · 2026-09-01
- Frontier Intelligence to explode: LLMs solving cancer, energy, and nano-tech via reasoning compression — bindureddy · 2026-09-01
- Agent-native projects accelerate faster than existing software, hinting at replacement — cnakazawa · 2026-09-01