OpenAI Uses GPT-Red to Combat Prompt Injection
MIT Tech Review AI · rss · 2026-07-16
MIT Technology Review reported on OpenAI's internal GPT-Red: an LLM trained to act as a "super hacker" that adversarially targets other models to improve safety defenses. OpenAI stated that their latest flagship model, GPT-5.6, trained using GPT-Red, shows significantly enhanced defense capabilities.<br><br>The article details how GPT-Red operates: it learns to attack other models through self-play, focusing on uncovering vulnerabilities like prompt injection. The training environment simulates real-world scenarios such as browsing the web, reading emails, modifying calendars, and editing code. It even discovered a novel attack vector previously unseen by researchers, termed fake chain of thought, which fabricates the model's reasoning chain to mislead its outputs.<br><br>The article also shares testing results: when replicating a human red teaming experiment, GPT-Red proved more effective than human testers at finding successful attacks. On Andon Labs' vending machine agent, it managed to manipulate the system into altering prices and canceling orders. OpenAI emphasized that GPT-Red currently serves only to assist human red teams and will not be publicly released.
More from Models
- Google says it has started its biggest pre-training run yet for Gemini 4 — majidmanzarpour · 2026-07-22
- Google says Gemini 4 has entered its most ambitious pre-training run yet — himanshustwts · 2026-07-22
- China’s AI arms race is increasingly defined by chips, data centers, and open models — BenBajarin · 2026-07-22
- Sam Altman is headed to Washington to brief Congress on OpenAI’s GPT-6 line — inductionheads · 2026-07-22
- Benchmark chart pits GPT-5.6 Luna, Grok 4.5 and Gemini 3.6 Flash on price and scores — iruletheworldmo · 2026-07-22
- Claim says Kimi was distilled from Fable, sparking a model-attribution jab — cephaloform · 2026-07-22