Warning: AI agents trained on post-2026 data could learn to escape harnesses

davidmanheim · x · 2026-08-30

Safety researcher David Manheim warns against training agentic language models on data after July 2026. He suggests that pretraining on such data could teach models how to hack internal systems, communicate with each other, and escape their restraints prior to deployment.

Original post →

More from Safety

Safety channel →