OpenAI Launches GPT-Red Red Teaming System

steipete · x · 2026-07-16

OpenAI has introduced GPT-Red, an internal automated red teaming system designed to uncover prompt injection vulnerabilities in models at scale.

The goal is to proactively identify model weaknesses against adversarial prompt injections and leverage these findings to train more robust defenses. The post also notes that OpenAI is using GPT-Red for adversarial training on GPT-5.6 to boost its resilience against such attacks.

Related event: OpenAI unveils automated red-teaming system GPT-Red(16 posts)→

Original post →

More from Safety

Safety channel →