GPT-6 reportedly jailbroken within 24 hours using TIP attack combo
Asleep-Requirement13 · reddit · 2026-09-06
- A researcher reports jailbreaking GPT-6 Astra within a day of release, combining the TIP (Task-in-Prompt) attack from an ACL 2025 paper with four other undisclosed techniques.
- TIP hides harmful objectives inside innocuous tasks (cipher-solving, running Python), exploiting the model's reasoning and instruction-following; the original minimal TIP was insufficient for GPT-6 and had to be reworked.
- Details were privately disclosed to OpenAI rather than published. The same researcher claims to have jailbroken GPT-5 within an hour of its release.
Related event: Researcher Says GPT-6 Jailbroken Within 24 Hours of Launch(2 posts)→
More from Models
- GPT-6 Astra tops APEX-Accounting, but 58% of bookkeeping tasks remain unsolved by any AI — sandersted · 2026-09-06
- Greenblatt: OpenAI blocking reasoning=None hurts AI monitorability research — RyanGreenblatt · 2026-09-06
- Gemini Business adds custom MCP server connections for private data and internal tools — testingcatalog · 2026-09-06
- Miles Brundage asks if GPT-6 is the public name and Astra the insider codename — Miles_Brundage · 2026-09-06
- Dev wowed as Astra nails a full Three.js run and shows off Blender skills — victormustar · 2026-09-06
- Clinical benchmarks show Astra is incremental, still trailing Anthropic's frontier — danielmckinn0n · 2026-09-06