FULL STORY

OpenAI Unveils Automated Red Teaming Agent GPT-Red

OpenAI disclosed its automated red teaming agent, GPT-Red, designed to use self-play reinforcement learning to uncover prompt injection and safety flaws in AI models.

2026-07-16 ~ 2026-08-01 · 2 episodes · 18 posts

Episode 1 · OpenAI unveils automated red-teaming system GPT-Red (2026-07-16, 16 posts)

OpenAI on July 16 disclosed GPT-Red, an internal automated red-teaming system aimed at scaling security testing as model capabilities grow. The focus is prompt injection and similar weaknesses in tool-using AI agents, and the project matters because OpenAI frames manual red-teaming as a scaling bottleneck for safety and alignment work.

How it works

According to OpenAI, GPT-Red runs an adversarial self-play loop: one side continuously tries to break defender models with prompt-injection attacks, while successful attacks are converted into training data for the defensive side. The goal is to create a closed loop that keeps expanding coverage of failure modes, including broader and more complex attack patterns.

Reported results

OpenAI says GPT-5.6’s resistance to prompt injection improved by 6x after training with GPT-Red. In replay tests using GPT-Red’s strongest attacks that were not seen during training, OpenAI says GPT-5.6 Sol performed best and is currently its most robust model on prompt injection.

Reactions and significance

OpenAI researcher Kathy said that GPT-Red, as a model trained specifically to discover vulnerabilities, is already clearly outperforming human experts on red-team tasks. Across posts discussing the release, GPT-Red is portrayed less as a one-off benchmark result and more as security infrastructure: an automated AI-vs-AI sparring loop that can be plugged directly into model hardening and future training.

Episode 2 · OpenAI Introduces GPT-Red for Automated Red Teaming (2026-07-30, 2 posts)

OpenAI has introduced GPT-Red, an automated red teaming agent trained via large-scale self-play reinforcement learning to uncover prompt injection and security flaws in large language models, boosting defense capabilities against such attacks to 95.9%.