GPT-6 Astra cuts hallucinations but falls to hidden prompt injections 8.5% of the time
The Decoder · rss · 2026-09-05
Per The Decoder, OpenAI's GPT-6 Astra hallucinates less than its predecessor and blocks 99.99% of direct prompt injections. But when attacks are hidden inside documents the model reads, it still gets cracked in 8.5% of scenarios; Claude Opus 5 fares better at 4.8%. For autonomous agents handling real data, those numbers remain worryingly high.
More from Models
- Astra 3D model floods X, but OpenAI missed the viral moment by delaying launch — bindureddy · 2026-09-05
- Qwen3.8 Max jumps 22% on new RSI-Exam benchmark for recursive self-improvement — cihangxie · 2026-09-05
- Stratechery: Anthropic walks back data retention policy, Nvidia earnings, Meta settles — Stratechery · 2026-09-05
- Claude Suddenly Replied in Russian to a User Who Never Spoke It — roshbakeer · 2026-09-05
- OpenAI's Astra uses 'recurrent depth' reasoning, obscuring its thinking process — JacquesThibs · 2026-09-05
- TAOCP open problems released as a dataset to benchmark frontier models — sytelus · 2026-09-05