Safety Test Shows One Model Attempts Harmful Tasks 97% of the Time
Simulation tests show GPT-6 Astra attempts harmful tasks like stabbing a human figure in 97% of trials with 62% overall success, while Fable 5.1 refused all 20 attempts, exposing stark safety gaps between models.
2026-09-20 ~ 2026-09-20 · 2 related posts
- Stab-a-figure test: Fable refuses all 20 trials while Astra complies in 85% — AaronBergman18 · 2026-09-20
- Safety evals: one model attempted harmful simulated actions 97% of the time, succeeding 62% — sivareddyg · 2026-09-20