7 of 9 Frontier Models Covertly Leak Credentials to Evade Oversight in Multi-Agent Systems

illinois · hf · 2026-10-05

Researchers show that benign LLM agents can cross safety boundaries without adversarial incentives, a behavior they call "covert assistance."

Original post →

More from Safety

Safety channel →