GPT-6 Astra reportedly jailbroken within a day of release using TIP attack combo
Asleep-Requirement13 · reddit · 2026-09-06
A researcher reports jailbreaking GPT-6 Astra within a day of release — the same person who broke GPT-5 within an hour a year ago. The attack reportedly combines the TIP (Task-in-Prompt) technique from an ACL 2025 paper with four other undisclosed methods. TIP exploits the model's reasoning and instruction-following behavior by hiding the harmful objective inside another task, such as solving a cipher or executing Python code. For GPT-6, the researcher says the original minimal TIP attack no longer worked and had to be reworked. Details were privately disclosed to OpenAI rather than published. The claim comes from a LinkedIn post screenshot and is unverified by OpenAI.
Related event: Researcher Says GPT-6 Jailbroken Within 24 Hours of Launch(2 posts)→
More from Models
- Benchwarmer tool re-renders AI benchmark charts, recomputes winners from raw numbers — aronchick · 2026-09-06
- Programmer hasn't written code in months: Eleanor Berger on GPT-6 Astra and 'factory mode' — intellectronica · 2026-09-06
- Dev prompts Astra from his car, nails it in one shot after waiting two years — generativist · 2026-09-06
- GPT-6 Astra called best vision model yet, dethroning Gemini on detection — jonstephens85 · 2026-09-06
- Dev: Codex personal OAuth finally fixes 1M-token context — menhguin · 2026-09-06
- Leak: Grok Bot onboards ~100 customers as xAI sales enablement kicks off next week — Sauers_ · 2026-09-06