GPT-6 Astra reportedly jailbroken within 24 hours using extended Task-in-Prompt attack

Asleep-Requirement13 · reddit · 2026-09-06

A researcher reports jailbreaking GPT-6 Astra within a day of release, combining the TIP (Task-in-Prompt) attack from an ACL 2025 paper with four other unnamed techniques. TIP exploits reasoning/instruction-following by hiding the harmful objective inside another task, such as solving a cipher or executing Python code. The researcher says the original minimal TIP attack no longer worked on GPT-6 and had to be reworked. Details were privately disclosed to OpenAI rather than published; the same researcher jailbroke GPT-5 within an hour of its release a year earlier.

Related event: Researcher claims GPT-6 Astra jailbroken within 24 hours of launch(5 posts)→

Original post →

More from Models

Models channel →