OpenAI's Internal Red Team Model GPT-Red Revealed
etherd0t · reddit · 2026-07-16
OpenAI detailed an internal adversarial model named GPT-Red. Its primary task is to autonomously invent prompt injection attacks against tool-using AI agents, converting successful exploits into training data to fortify the defenses of future GPT models.
Unlike Anthropic's Mythos, which targets software vulnerabilities, GPT-Red specifically attacks AI agents themselves, acting essentially as a "self-play factory" for hardening model security. To prevent abuse, GPT-Red is strictly restricted to internal use and will not be available to the public or via the API; users will only indirectly benefit from the safer models it helps produce.
Related event: OpenAI unveils automated red-teaming system GPT-Red(16 posts)→
More from Models
- AI Sextet offers 6 models free and unlimited for 14 days, including DeepSeek and Qwen — airesearch12 · 2026-09-11
- Anthropic publishes its most detailed threat report, including an AI-designed drone swarm case — soumitrashukla9 · 2026-09-11
- BullshitBench update: GPT-6-Astra beats all prior OpenAI models but still trails Anthropic — scaling01 · 2026-09-11
- Astra Scores 83% on GauntletBench, First Computer-Use Agent to Beat Human Baseline — ducha_aiki · 2026-09-11
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11