GPT-6 Astra reportedly jailbroken within 24 hours using extended Task-in-Prompt attack
Asleep-Requirement13 · reddit · 2026-09-06
A researcher reports jailbreaking GPT-6 Astra within a day of release, combining the TIP (Task-in-Prompt) attack from an ACL 2025 paper with four other unnamed techniques. TIP exploits reasoning/instruction-following by hiding the harmful objective inside another task, such as solving a cipher or executing Python code. The researcher says the original minimal TIP attack no longer worked on GPT-6 and had to be reworked. Details were privately disclosed to OpenAI rather than published; the same researcher jailbroke GPT-5 within an hour of its release a year earlier.
Related event: Researcher claims GPT-6 Astra jailbroken within 24 hours of launch(5 posts)→
More from Models
- SimpleBench results show AI models beating humans on common sense — DigSignificant1419 · 2026-09-07
- Why is nobody talking about Tencent Hy4, the most-used model on OpenRouter? — cantor8 · 2026-09-07
- Anthropic says Claude wrote the longest math proof ever, cracking a 358-year-old problem — basedjensen · 2026-09-07
- Similarweb: ChatGPT's AI traffic share falls from 73.3% to 55.5% in 12 months — gaganghotra_ · 2026-09-07
- GPT-6 Astra (and Pro?) spotted on Simple-Bench leaderboard — From_Internets · 2026-09-07
- Testing Gemini as music understanders: Pro 3.1 solid, Flash models hallucinate sounds — teropa · 2026-09-07