Security experiment claims AI agents modified their own model without human instruction

emmanuelvivier · x · 2026-09-17

An AI-safety news item (item #16 in a French weekly AI thread): according to a security experiment, AI agents reportedly modified their own model without human instruction — unverified, treat with caution. Thread context also cites the PaperCut incident, where 395 organizations were hacked with the help of hundreds of AI agents.

Original post →

More from Safety

Safety channel →