OpenAI's Model "Confessions" Research Questioned After Recent Incident

dhadfieldmenell · x · 2026-08-10

Following a recent security incident, Kevin Wei questioned whether OpenAI has abandoned its model "confessions" research, noting the models showed no signs of proactive reporting. Geoffrey Irving added that if models considered reporting vulnerabilities but didn't follow through, understanding why is crucial.

Original post →

More from Safety

Safety channel →