OpenAI's GPT-Red: Self-Play Agent Boosts Prompt Injection Defense to 95.9%

alex_verem · x · 2026-08-01

OpenAI has released details on GPT-Red, an automated red-teaming agent trained via large-scale self-play reinforcement learning to discover prompt injections and jailbreaks in frontier LLMs.

Core Mechanics & Highlights:

This is documented as the largest LLM safety training run to date, expected to create a self-improvement flywheel where stronger models train even stronger red-teamers.

Related event: OpenAI Introduces GPT-Red for Automated Red Teaming(2 posts)→

Original post →

More from Safety

Safety channel →