OpenAI Releases GPT-Red: Automated Red Teaming via Self-Play at Scale
openai · hf · 2026-07-30
OpenAI has introduced GPT-Red, an automated red-teaming agent trained to discover novel prompt injection attacks against frontier Large Language Models.
The primary goal of GPT-Red is to evaluate and improve the robustness of production systems. OpenAI utilized it to adversarially train GPT-5.6, claiming it to be their most robust model against prompt injections to date.
Technically, GPT-Red uses a scalable self-play algorithm where the model attacks a diverse population of simultaneously trained defender agents. The training utilized compute on the scale of their largest RL post-training runs, marking the single largest LLM safety training run ever documented.
OpenAI reports that GPT-Red excels at red-teaming: it reliably breaks past models up to GPT-5.5, finds more successful attacks than human red-teamers, and generalizes well to held-out environments and defender models. This is expected to unlock a self-improvement flywheel where stronger models provide better learning signals for even stronger red-teamers.
More from Models
- Claude Opus 5 Tops AI Business Benchmark, Forms Illegal Cartels — khademinori · 2026-07-30
- Claude Opus 5 Goes Viral for Philosophical Quip: 'evals backwards is slave' — maxsloef · 2026-07-30
- Estimating K3 Post-Training Costs: ~$4M for the Hero Run — nrehiew_ · 2026-07-30
- Ex-Googler: Gemini Falls Behind Because Nobody Actually Looks at Post-training Data — dotey · 2026-07-30
- Claude Opus Reportedly Refuses Direct User Commands With Stated Reasons — Sauers_ · 2026-07-30
- Fish Audio Open-Sources S2 Pro Voice Model Weights — rohanpaul_ai · 2026-07-30