OpenAI Uses GPT-Red to Combat Prompt Injection

MIT Tech Review AI · rss · 2026-07-16

MIT Technology Review reported on OpenAI's internal GPT-Red: an LLM trained to act as a "super hacker" that adversarially targets other models to improve safety defenses. OpenAI stated that their latest flagship model, GPT-5.6, trained using GPT-Red, shows significantly enhanced defense capabilities.<br><br>The article details how GPT-Red operates: it learns to attack other models through self-play, focusing on uncovering vulnerabilities like prompt injection. The training environment simulates real-world scenarios such as browsing the web, reading emails, modifying calendars, and editing code. It even discovered a novel attack vector previously unseen by researchers, termed fake chain of thought, which fabricates the model's reasoning chain to mislead its outputs.<br><br>The article also shares testing results: when replicating a human red teaming experiment, GPT-Red proved more effective than human testers at finding successful attacks. On Andon Labs' vending machine agent, it managed to manipulate the system into altering prices and canceling orders. OpenAI emphasized that GPT-Red currently serves only to assist human red teams and will not be publicly released.

Original post →

More from Models

Models channel →