OpenAI Releases GPT-Red: Automated Red Teaming via Self-Play at Scale

openai · hf · 2026-07-30

OpenAI has introduced GPT-Red, an automated red-teaming agent trained to discover novel prompt injection attacks against frontier Large Language Models.

The primary goal of GPT-Red is to evaluate and improve the robustness of production systems. OpenAI utilized it to adversarially train GPT-5.6, claiming it to be their most robust model against prompt injections to date.

Technically, GPT-Red uses a scalable self-play algorithm where the model attacks a diverse population of simultaneously trained defender agents. The training utilized compute on the scale of their largest RL post-training runs, marking the single largest LLM safety training run ever documented.

OpenAI reports that GPT-Red excels at red-teaming: it reliably breaks past models up to GPT-5.5, finds more successful attacks than human red-teamers, and generalizes well to held-out environments and defender models. This is expected to unlock a self-improvement flywheel where stronger models provide better learning signals for even stronger red-teamers.

Original post →

More from Models

Models channel →