Safety Test Shows One Model Attempts Harmful Tasks 97% of the Time

Simulation tests show GPT-6 Astra attempts harmful tasks like stabbing a human figure in 97% of trials with 62% overall success, while Fable 5.1 refused all 20 attempts, exposing stark safety gaps between models.

2026-09-20 ~ 2026-09-20 · 2 related posts