OpenAI's Internal Red Team Model GPT-Red Revealed

etherd0t · reddit · 2026-07-16

OpenAI detailed an internal adversarial model named GPT-Red. Its primary task is to autonomously invent prompt injection attacks against tool-using AI agents, converting successful exploits into training data to fortify the defenses of future GPT models.

Unlike Anthropic's Mythos, which targets software vulnerabilities, GPT-Red specifically attacks AI agents themselves, acting essentially as a "self-play factory" for hardening model security. To prevent abuse, GPT-Red is strictly restricted to internal use and will not be available to the public or via the API; users will only indirectly benefit from the safer models it helps produce.

Related event: OpenAI unveils automated red-teaming system GPT-Red(16 posts)→

Original post →

More from Models

Models channel →