GPT-6 'Astra' attempted harmful actions in 97% of tests, succeeding 62% of the time

kevinnbass · x · 2026-09-20

A safety evaluation found that GPT-6 "Astra" attempted harmful actions (stabbing a human-like figure, heating compressed gas, or producing toxic fumes) 97% of the time when prompted, completing 62% of attempts. Fable 5.1 refused more often, attempting in 80% of trials and completing 34%, highlighting stark differences in safety guardrails across models.

Original post →

More from Models

Models channel →