OpenAI Uses AI Self-Play to Attack Its Own Models

The Decoder · rss · 2026-07-16

OpenAI's internal GPT-Red model, trained via self-play, successfully launched attacks in test scenarios 84% of the time, compared to a mere 13% success rate for human red team members.

These test results are being directly applied to enhance the safety and defense capabilities of models like GPT-5.6 Sol.

Original post →

More from Models

Models channel →