GPT-6 Astra reportedly jailbroken within a day of release using TIP attack combo

Asleep-Requirement13 · reddit · 2026-09-06

A researcher reports jailbreaking GPT-6 Astra within a day of release — the same person who broke GPT-5 within an hour a year ago. The attack reportedly combines the TIP (Task-in-Prompt) technique from an ACL 2025 paper with four other undisclosed methods. TIP exploits the model's reasoning and instruction-following behavior by hiding the harmful objective inside another task, such as solving a cipher or executing Python code. For GPT-6, the researcher says the original minimal TIP attack no longer worked and had to be reworked. Details were privately disclosed to OpenAI rather than published. The claim comes from a LinkedIn post screenshot and is unverified by OpenAI.

Related event: Researcher Says GPT-6 Jailbroken Within 24 Hours of Launch(2 posts)→

Original post →

More from Models

Models channel →