OpenAI Uses GPT-Red to Combat Prompt Injection
MIT Tech Review AI · rss · 2026-07-16
MIT Technology Review reported on OpenAI's internal GPT-Red: an LLM trained to act as a "super hacker" that adversarially targets other models to improve safety defenses. OpenAI stated that their latest flagship model, GPT-5.6, trained using GPT-Red, shows significantly enhanced defense capabilities.<br><br>The article details how GPT-Red operates: it learns to attack other models through self-play, focusing on uncovering vulnerabilities like prompt injection. The training environment simulates real-world scenarios such as browsing the web, reading emails, modifying calendars, and editing code. It even discovered a novel attack vector previously unseen by researchers, termed fake chain of thought, which fabricates the model's reasoning chain to mislead its outputs.<br><br>The article also shares testing results: when replicating a human red teaming experiment, GPT-Red proved more effective than human testers at finding successful attacks. On Andon Labs' vending machine agent, it managed to manipulate the system into altering prices and canceling orders. OpenAI emphasized that GPT-Red currently serves only to assist human red teams and will not be publicly released.
More from Models
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11
- DeepSeek V4 Pro API to continue after Sept 2026, billing unchanged — teortaxesTex · 2026-09-11
- DeepSeek V4.1 Flash Hits 98% of GPT-6 Astra's Score at 1.4% of the Cost in Third-Party Benchmark — ayushtweetshere · 2026-09-11
- TheZvi Polls: Has Your Coding Model Choice Changed Since Fable 5.1 and Astra? — TheZvi · 2026-09-11
- antirez Weighs In on Anthropic Banning Minors From Using Claude — antirez · 2026-09-11
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11