GPT-6 Astra cuts hallucinations but falls to hidden prompt injections 8.5% of the time

The Decoder · rss · 2026-09-05

Per The Decoder, OpenAI's GPT-6 Astra hallucinates less than its predecessor and blocks 99.99% of direct prompt injections. But when attacks are hidden inside documents the model reads, it still gets cracked in 8.5% of scenarios; Claude Opus 5 fares better at 4.8%. For autonomous agents handling real data, those numbers remain worryingly high.

Original post →

More from Models

Models channel →